Season one · stop 3 · Attention

Letting the Translator Look Back

“Here's the idea that made modern models possible. One translator has to squeeze a whole sentence into a single note. The other is allowed to look back. Watch what happens as sentences get longer.”

Unpacking the world: 1.9 MB of drawings and measured data.

AgenticAmit A field guide to attention

Field guide · attention

Letting the translator look back.

Before 2014, a translation model had to squeeze a whole sentence into one short note and translate from memory. Attention let it look back at the words while it wrote. We trained both kinds and watched.

Scroll to cross the river ↓

The task · a toy language pair we madereal French words
Translate this.
English
French
Pinned · the one-note designNIPS 2014
Sutskever, Vinyals and Le: Our method uses a multilayered Long Short-Term Memory (LSTM) to map the input sequence to a vector of a fixed dimensionality, and then another deep LSTM to decode the target sequence from the vector.

Sutskever, Vinyals & Le · Sequence to Sequence Learning with Neural Networks

Crossing 1 · Englishread in
The note
Everything, in one vector.

Crossing 1 · Frenchwritten from the note
Crossing 2 · Englishkept, word by word
Crossing 2 · Frencheach word looks back
Pinned · the ideaICLR 2015
Bahdanau et al.: In this paper, we conjecture that the use of a fixed-length vector is a bottleneck in improving the performance of this basic encoder–decoder architecture, and propose to extend this by allowing a model to automatically (soft-)search for parts of a source sentence that are relevant to predicting a target word, without having to form these parts as a hard segment explicitly.

Bahdanau, Cho & Bengio · Neural Machine Translation by Jointly Learning to Align and Translate

Pinned · the problem, measuredSSST-8 2014
Cho et al.: We show that the neural machine translation performs relatively well on short sentences without unknown words, but its performance degrades rapidly as the length of the sentence and the number of unknown words increase.

Cho, van Merriënboer, Bahdanau & Bengio · On the Properties of Neural Machine Translation

Fig. 03 · where each French word lookedreal attention weights
The alignment grid

Fig. 04 · our two translators
Longer sentences, one note: it breaks.
one note (no attention)with attention

Pinned · the same shape, on real FrenchICLR 2015
Bahdanau et al. Figure 2: BLEU score against sentence length for RNNsearch-50, RNNsearch-30, RNNenc-50 and RNNenc-30. Figure 2: The BLEU scores of the generated translations on the test set with respect to the lengths of the sentences. The results are on the full test set which includes sentences having unknown words to the models.

Bahdanau, Cho & Bengio · Neural Machine Translation by Jointly Learning to Align and Translate

Pinned · their gridICLR 2015
Bahdanau et al. Figure 3(a): alignment between an English sentence about the European Economic Area and its French translation, showing zone économique européenne aligned in reverse order to European Economic Area. Bahdanau et al. Figure 3 caption.

Bahdanau, Cho & Bengio · Figure 3

The catch · articles, over of themcrossing 2
Pinned · the argumentNAACL 2019 · EMNLP 2019
Jain and Wallace: We find that they largely do not. For example, learned attention weights are frequently uncorrelated with gradient-based measures of feature importance, and one can identify very different attention distributions that nonetheless yield equivalent predictions. Wiegreffe and Pinter: A recent paper claims that Attention is not Explanation (Jain and Wallace, 2019). We challenge many of the assumptions underlying this work, arguing that such a claim depends on one's definition of explanation, and that testing it needs to take into account all elements of the model.

Jain & Wallace · Attention is not Explanation · Wiegreffe & Pinter · Attention is not not Explanation

Fig. 05 · every word looks at every wordlooks per sentence
The grid grows with the square.

Pinned · where it wentNIPS 2017
Vaswani et al.: We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.

Vaswani et al. · Attention Is All You Need

no belt at all!

Take this with you

Don't memorise the sentence. Keep it, and look back.

One fixed note can only hold so much, so long sentences fall apart. Attention keeps every word within reach and decides, at each step, where to look. That one idea, taken all the way, became the Transformer.

Amit, handing something over