1. Switch
  2. Compute
  3. Deduce
  4. Probability
  5. Information
  6. Vectors
  7. Derivatives
  8. Optimize
  9. Neurons
  10. Generalize
  11. Attention
  12. LLM

Chapter 11 · Language models

Predicting the next word, at planetary scale

An LLM is a Transformer trained to minimize Shannon's cross-entropy over trillions of words. This chapter brings every piece together: how it generates text, why it improves predictably as it grows, how it becomes an assistant like ChatGPT, and why it sometimes makes things up.

On 30 November 2022 OpenAI opened an experimental chat to the public. Within five days it had a million users and, by some estimates, it became the fastest-growing consumer application in history up to that point. Under the hood, ChatGPT was a model of the GPT family, that is, a Transformer trained to do a single thing: predict the next token. Everything we have seen in the previous chapters happens in it at once.

The recipe

  1. Tokenize. Text is cut into frequent units with Byte Pair Encoding, a 1994 compression algorithm (Gage) that Sennrich and colleagues adapted to language in 2016. Common words are a single token; rare ones are split into pieces.
  2. Represent. Each token becomes an embedding, and dozens of attention blocks transform it into a vector that summarizes the whole preceding context.
  3. Predict. A final matrix turns that vector into a score (logit) zi for each of the ~100,000 tokens in the vocabulary, and a softmax turns the scores into a distribution.
  4. Train. The cross-entropy (the negative log-likelihood) is minimized over trillions of tokens, with backpropagation and AdamW, for weeks and on thousands of GPUs.
ℒ(θ)=−1T∑t=1Tlog⁡pθ(wt∣w1,…,wt−1).

Nothing in this formula mentions grammar, facts or reasoning. But predicting well the text that people write requires, to some extent, all of that. Ilya Sutskever put it this way: compressing the data well forces you to learn what generated it.

Generating: temperature and sampling

To write, the model computes the distribution of the next token, picks one, appends it to the text and repeats. How it picks matters a great deal. The best-known parameter is the temperature T:

pi=ezi/T∑jezj/T.

With T→0 the most probable token always wins (greedy decoding, repetitive). With T=1 the learned distribution is sampled. With large T everything tends towards equally likely (creative… and then incoherent). The name is no accident: it is the Boltzmann distribution of statistical physics, and it has a characterization that connects back to Shannon.

Theorem (softmax as maximum entropy; Jaynes, 1957)

Among all distributions p over the tokens with a fixed mean logit, ∑ipizi=c, the one with maximum entropy is pi∝ezi/T, for the temperature T that satisfies the constraint.

Proof

Let qi=ezi/T/Z. For any p that satisfies the constraint, by Gibbs' inequality,

0≤DKL(p‖q)=−H(p)−∑ipilog⁡qi=−H(p)−cT+log⁡Z.

Hence H(p)≤log⁡Z−c/T=H(q), with equality only if p=q (all in nats).

In practice it is combined with truncation: top-k keeps the k most probable tokens, and top-p (nucleus sampling, Holtzman et al., 2020) keeps the smallest set whose cumulative probability exceeds p. This avoids the oddities of the tail of the distribution without losing variety.

The distribution of the next token for “The cat is sitting on the”. Lower the temperature so that “mat” almost always wins; raise it to give “keyboard” a chance. Tokens cut by top-k or top-p are dimmed and their probability is shared among the rest. With “Sample 100” you can compare the frequencies obtained with the probabilities.

Scaling laws

In January 2020 Jared Kaplan and his colleagues at OpenAI found an astonishing regularity: the loss of a language model falls as a power law of the number of parameters N, the number of training tokens D and the compute used, across seven orders of magnitude. In 2022 Jordan Hoffmann's team at DeepMind fitted an explicit form after training more than 400 models:

L(N,D)=E⏟irreducible+ANα+BDβ,C≈6NDoperations.

The term E is the loss no model can remove: the entropy of text itself, the uncertainty that remains even with perfect knowledge of the language. The other two are the price of a finite model and of finite data. With a fixed compute budget C, minimizing L is an optimization problem with a simple solution.

Result (compute-optimal allocation; Hoffmann et al., 2022)

Minimizing L(N,D) subject to 6ND=C gives N∗∝Ca and D∗∝Cb, with a=βα+β and b=αα+β. Empirically a≈b≈0.5: parameters and data should grow at the same pace, at about 20 tokens per parameter.

Iso-compute curves: for a fixed budget C, each point is a different split between model size and data (D=C/6N). A model that is too small does not make use of the data; one that is too large does not see enough of it. GPT-3 (175 billion parameters, 300 billion tokens) was oversized: Chinchilla, with 70 billion parameters and 1.4 trillion tokens and similar compute, beat it. Parameters from the replication by Besiroglu et al. (2024); loss in nats per token.

Scaling laws turned AI into an engineering discipline with predictable budgets. They explain the race for data centres and also its limits: the improvement is slow (doubling compute cuts the reducible part of the loss by only a few per cent) and high-quality text data is not infinite.

From predictor to assistant

A pretrained model continues texts but does not follow instructions: ask it something and it may reply with another question, because that is how many texts on the internet go on. To turn it into an assistant it is fine-tuned on example conversations and, above all, with reinforcement learning from human feedback (RLHF; Christiano et al., 2017; Ouyang et al., 2022). This is where the last branch of the table comes in: game theory and reinforcement learning, which from the Bellman equation (1957) to AlphaGo (2016) study agents that learn from rewards.

Step 1: a reward model. People compare pairs of answers. A function r(x,y) is fitted with the Bradley–Terry model (1952), originally designed to rank players from the games they played:

P(yA≻yB∣x)=σ(r(x,yA)−r(x,yB)).

Step 2: optimize the policy without straying too far. We look for a model π that earns a lot of reward but still talks like the original model πref, penalizing the KL divergence:

maxπ⁡𝔼y∼π(⋅∣x)[r(x,y)]−βDKL(π(⋅∣x)‖πref(⋅∣x)).
Theorem (optimal policy with KL regularization)

The maximum is attained at

π∗(y∣x)=1Z(x)πref(y∣x)exp⁡(r(x,y)β).
Proof

Dividing the objective by −β and rearranging, maximizing it is equivalent to minimizing

∑yπ(y)log⁡π(y)πref(y)er(y)/β=DKL(π‖π∗)−log⁡Z,

which by Gibbs' inequality is smallest exactly when π=π∗.

It is a softmax again: the final model reweights the probabilities of the original one by the exponential of the reward, and β plays the role of the temperature. This result is the basis of DPO (Rafailov et al., 2023), which solves the formula for r and trains directly on the preferences, with no explicit reinforcement learning. Since 2024, reasoning models have taken the idea further: they are trained with reinforcement on problems with verifiable answers (mathematics, code) and learn to “think” through long chains of steps before answering, trading compute at answer time for accuracy.

Why do they hallucinate?

An LLM sometimes states false things with total confidence. Part of the problem has a statistical explanation. A model trained by maximum likelihood tends to be calibrated: when it assigns 70% to something, it is right about 70% of the time. Kalai and Vempala (2024) proved that this comes at a cost.

Theorem (calibrated models hallucinate; Kalai and Vempala, 2024, informal)

For arbitrary facts that follow no pattern (such as the birthday of a little-known person), a calibrated language model must generate false facts at a rate roughly at least equal to the fraction of such facts that appear exactly once in the training data. That fraction is the Good–Turing estimate of the “unseen mass”.

What has been seen only once cannot be told apart well from what has never been seen. Reducing hallucinations takes more than predicting well: teaching the model to abstain, to consult external sources or to verify. No amount of scaling the training data removes this tension on its own.

So, why does AI work?

Let us walk the chain one last time. Turing showed that a universal machine can run any procedure, including a learned one. Probability gave the language for learning from uncertain data, and the chain rule turned text into a sequence of predictions. Shannon turned “predicting well” into a number, cross-entropy, which is at once likelihood and compression. Linear algebra provided a space where meaning is geometry; calculus and backpropagation, the gradient of billions of parameters for the price of a few evaluations; and stochastic optimization, a cheap way of walking down it. Neural networks contributed functions able to approximate almost anything, and attention, the right inductive bias for language and for hardware. At scale, the loss falls according to power laws, and reinforcement with KL regularization turns the predictor into an assistant.

And open questions remain, very mathematical ones: why networks that could memorize everything generalize, what exactly they represent inside, and how to guarantee that they do what we want. That these questions are still open is perhaps the best reason to study the mathematics of AI.

Review the full history →Back to the contents

References

  1. E. T. Jaynes (1957). “Information Theory and Statistical Mechanics”. Physical Review, 106(4).
  2. R. A. Bradley and M. E. Terry (1952). “Rank Analysis of Incomplete Block Designs”. Biometrika, 39(3/4).
  3. R. Sennrich, B. Haddow and A. Birch (2016). “Neural Machine Translation of Rare Words with Subword Units”. ACL.
  4. P. Christiano et al. (2017). “Deep Reinforcement Learning from Human Preferences”. NeurIPS.
  5. A. Radford et al. (2019). “Language Models are Unsupervised Multitask Learners”. OpenAI.
  6. J. Kaplan et al. (2020). “Scaling Laws for Neural Language Models”. arXiv:2001.08361.
  7. T. Brown et al. (2020). “Language Models are Few-Shot Learners”. NeurIPS.
  8. A. Holtzman et al. (2020). “The Curious Case of Neural Text Degeneration”. ICLR.
  9. J. Hoffmann et al. (2022). “Training Compute-Optimal Large Language Models”. NeurIPS.
  10. L. Ouyang et al. (2022). “Training language models to follow instructions with human feedback”. NeurIPS.
  11. R. Rafailov et al. (2023). “Direct Preference Optimization: Your Language Model is Secretly a Reward Model”. NeurIPS.
  12. R. Schaeffer, B. Miranda and S. Koyejo (2023). “Are Emergent Abilities of Large Language Models a Mirage?”. NeurIPS.
  13. A. T. Kalai and S. S. Vempala (2024). “Calibrated Language Models Must Hallucinate”. STOC.
  14. T. Besiroglu, E. Erdil, M. Barnett and J. You (2024). “Chinchilla Scaling: A replication attempt”. arXiv:2404.10102.