Chapter 10 · Transformers
Attention: every word looks at all the others
In 2017 eight researchers at Google published a paper with a provocative title: “Attention Is All You Need”. Their architecture, the Transformer, replaced word-by-word reading with an operation in which each word decides, with a dot product and a softmax, which others to pay attention to.
In this chapter
Read this sentence: “The cats that chased the mouse were very tired”. To know that the verb must be plural you have to remember “cats”, five words back, and not be misled by “mouse”, which is closer. An -gram model that only looks at the previous two or three words cannot do it. Capturing long-range dependencies was for decades the great problem of language modelling.
Before attention: reading left to right
Recurrent networks (Elman, 1990) read text word by word, updating a state vector, a kind of summarized memory: . In practice they forgot quickly, because of the vanishing gradient along the chain. The LSTMs of Sepp Hochreiter and Jürgen Schmidhuber (1997) added gates that decide what to keep and what to forget, and they dominated machine translation and speech recognition for twenty years.
But they were still sequential. To translate, an encoder compressed the whole input sentence into a single vector, and a decoder generated the translation from it (Sutskever, Vinyals and Le, 2014). A long sentence does not fit well in one vector. In 2014 Dzmitry Bahdanau, Kyunghyun Cho and Yoshua Bengio proposed letting the decoder, at each step, look at every word of the input and choose which ones to focus on. They called it attention.
Attention: a soft lookup in a dictionary
Think of a Python dictionary: you look up a key and get its value. Attention is a “soft”, differentiable version of that lookup. Each word (each token) produces three vectors from its embedding , through three learned matrices:
- a query : “what am I looking for?”;
- a key : “what do I offer?”;
- a value : “here is my information if you pick me”.
The match between the query of word and the key of word is a dot product. A softmax turns those scores into weights that add up to 1, and the output is the average of the values weighted by those weights. With all the words at once, in matrix form:
Look at the agreement head: “were” attends to “cats” and not to “mouse”, even though “mouse” is closer. Distance does not matter. Any word can query any other word directly, in a single step.
Why divide by ?
The factor in the formula is not decorative. It comes from a variance calculation.
If the components of are independent, with mean 0 and variance 1, then
Proof
. Each term has mean and variance , and the terms are independent, so the variances add up.
With , the scores would have a standard deviation above 11. The softmax of such disparate numbers is practically “one and all zeros”, and in that regime its gradient is almost zero: the network would stop learning. Dividing by brings the variance back to 1. You can check it in the figure by switching the division off.
The Transformer block
A Transformer stacks dozens of identical blocks. Each has two parts:
- Multi-head attention. Several attentions in parallel, each with its own , which can specialize in different relations (syntax, position, coreference…). Their outputs are concatenated and mixed with another matrix.
- A two-layer neural network applied to each position separately. Much of the model's factual “knowledge” is believed to be stored there.
Around each part there is a residual connection, , and a normalization. Attention alone does not distinguish word order: “the dog bites the man” and “the man bites the dog” would look the same. That is why a positional encoding is added to each embedding. The original paper used sines and cosines of different frequencies,
chosen because shifting the position amounts to a rotation, a linear transformation. The “previous word” head in the figure relies on exactly that property.
Generating text: the causal mask
A model like GPT is a “decoder-only” Transformer: at each position it predicts the next token. To stop it from cheating during training, each position is prevented from looking at later ones by putting in those scores before the softmax. Thanks to this causal mask, a single pass over a text of tokens yields training predictions at once.
Here lies the practical reason for the Transformer's success: unlike recurrent networks, it processes every position in parallel, with matrix multiplications that GPUs execute very fast. It has a price: the attention matrix has entries, so the cost grows with the square of the context length. Much current research (FlashAttention, sparse attention, state-space models) tries to make that term cheaper.
What we know in theory
Transformers with positional encoding can approximate to arbitrary precision any continuous function from fixed-length sequences to sequences, on a compact domain.
It is the sequence version of the universal approximation theorem. Pérez, Barceló and Marinković (2021) also proved that, with unlimited arithmetic precision, Transformers are Turing-complete: the circle closes with the first chapter. Even more intriguing is in-context learning: it has been shown that attention can implement, within a single pass, algorithms such as a step of gradient descent on examples given in the text itself (von Oswald et al., 2023). It is a clue to how an LLM “learns” from the examples in a prompt without changing its weights.
We now have all the pieces. In the last chapter we put them together at scale: trillions of tokens, hundreds of billions of parameters and an objective as simple as predicting the next word.
References
- J. L. Elman (1990). “Finding Structure in Time”. Cognitive Science, 14(2).
- S. Hochreiter and J. Schmidhuber (1997). “Long Short-Term Memory”. Neural Computation, 9(8).
- I. Sutskever, O. Vinyals and Q. V. Le (2014). “Sequence to Sequence Learning with Neural Networks”. NeurIPS.
- D. Bahdanau, K. Cho and Y. Bengio (2015). “Neural Machine Translation by Jointly Learning to Align and Translate”. ICLR (arXiv 2014).
- A. Vaswani et al. (2017). “Attention Is All You Need”. NeurIPS.
- C. Yun, S. Bhojanapalli, A. S. Rawat, S. J. Reddi and S. Kumar (2020). “Are Transformers universal approximators of sequence-to-sequence functions?”. ICLR.
- J. Pérez, P. Barceló and J. Marinković (2021). “Attention is Turing Complete”. JMLR, 22.
- J. Su et al. (2021). “RoFormer: Enhanced Transformer with Rotary Position Embedding”. arXiv:2104.09864.
- J. von Oswald et al. (2023). “Transformers Learn In-Context by Gradient Descent”. ICML.