1. Switch
  2. Compute
  3. Deduce
  4. Probability
  5. Information
  6. Vectors
  7. Derivatives
  8. Optimize
  9. Neurons
  10. Generalize
  11. Attention
  12. LLM

Chapter 03 · Probability

Learning from data

Probability puts numbers on uncertainty. Two eighteenth-century results, the law of large numbers and Bayes' theorem, explain why more data brings us closer to the truth and how to update what we believe as each new data point arrives.

In 1654 the Chevalier de Méré put a gambling problem to Blaise Pascal: if a game of several rounds is interrupted before it ends, how should the stakes be fairly divided? Pascal wrote about it to Pierre de Fermat, and the calculus of probability was born in that correspondence. Three centuries later, the same mathematics that split wagers decides which word ChatGPT writes next.

A language model does not “know” which word comes next: it assigns a probability to every possible word. Everything it does rests on this language.

Random variables and distributions

A random variable X is a quantity whose value depends on chance: the outcome of a coin, the height of a randomly chosen person, or the next word of a text. Its distribution p(x)=P(X=x) assigns each outcome a number between 0 and 1, and they all add up to 1:

p(x)≥0,∑xp(x)=1.

The expected value is the average weighted by the probabilities, 𝔼[X]=∑xxp(x), and the variance measures how much the variable spreads around that mean, Var⁡(X)=𝔼[(X−𝔼[X])2].

The law of large numbers

If a coin lands heads with a probability p we do not know, intuition says that after many flips the proportion of heads will approach p. Jakob Bernoulli spent twenty years proving it, and the result was published in 1713, eight years after his death, in Ars Conjectandi. He called it his “golden theorem”.

Theorem (weak law of large numbers)

Let X1,X2,… be independent variables with the same distribution, mean μ and finite variance σ2, and let X¯n=1n∑i=1nXi be their sample mean. For every ε>0,

limn→∞⁡P(|X¯n−μ|>ε)=0.
Proof (with Chebyshev's inequality)

By independence, Var⁡(X¯n)=σ2/n. Chebyshev's inequality says that a variable strays from its mean by more than ε with probability at most Var⁡/ε2, so

P(|X¯n−μ|>ε)≤σ2nε2⟶0as n→∞.

The proof gives more than the limit: the typical error of the mean shrinks like σ/n. To be twice as precise you need four times as much data. This law of diminishing returns reappears, with other exponents, in the scaling laws of LLMs.

A biased coin with a hidden probability of heads p. Left: the frequency of heads approaches p as n grows (law of large numbers; the horizontal axis is logarithmic). Right: what we believe about p after each flip, according to Bayes' theorem. Try starting from the belief that the coin is fair and see how long it takes to convince you otherwise.

Bayes' theorem

The law of large numbers speaks about what happens in the long run, but we almost never have infinite data. What can we say about p after only ten flips? The Reverend Thomas Bayes tackled this inverse problem in an essay his friend Richard Price published in 1763, after his death. Laplace rediscovered and generalized it in 1774.

Theorem (Bayes)

For a hypothesis H and data D with P(D)>0,

P(H∣D)⏟posterior=P(D∣H)⏞likelihoodP(H)⏞priorP(D).
Proof

By the definition of conditional probability, P(H∩D)=P(H∣D)P(D)=P(D∣H)P(H). Just solve for P(H∣D).

The theorem fits on one line, but its reading is deep: the prior summarizes what we believed before seeing the data, the likelihood measures how well each hypothesis explains what was observed, and the posterior is the updated belief. Each new data point turns today's posterior into tomorrow's prior.

For the coin, if the prior over p is a Beta(a,b) distribution and we observe k heads in n flips, the posterior is again a Beta:

p∣data∼Beta(a+k,b+n−k).

With the uniform prior (a=b=1), the probability that the next flip lands heads is k+1n+2. This is Laplace's rule of succession, with which in 1814 he computed the probability that the sun will rise tomorrow given that it has risen every day of recorded history. It was the first “smoothing” of an estimate, and statistical language models would use it two hundred years later so as not to assign zero probability to words they had never seen.

Maximum likelihood: the loss function of AI

Between 1912 and 1922 Ronald Fisher proposed a more direct criterion: among all possible values of the parameter, keep the one that makes the observed data most probable. If the data are independent the likelihood is a product, and since multiplying many small numbers causes numerical trouble, its logarithm is maximized instead:

θ^=argmaxθ⁡∏i=1np(xi∣θ)=argminθ⁡(−∑i=1nlog⁡p(xi∣θ))⏟negative log-likelihood.

For the coin we get the expected p^=k/n. But the expression on the right deserves a second look: it is exactly the function minimized when training any language model. In the next chapter we will see that it also has a name in information theory: cross-entropy.

The key idea

Training a model means choosing the parameters that make the training data most probable. Everything else (networks, gradients, attention) is a way of expressing and optimizing that probability.

The chain rule of probability

One piece is missing. How do you assign a probability to a whole sentence, when the number of possible sentences is astronomical? With an elementary identity that follows from applying the definition of conditional probability over and over:

Chain rule
P(w1,w2,…,wT)=∏t=1TP(wt∣w1,…,wt−1).

The probability of a text is the product of the probabilities of each word conditioned on all the previous ones. This identity turns the impossible problem of modelling whole texts into a manageable one: predicting the next word. An LLM is a gigantic approximator of P(wt∣w<t), and it generates text by sampling from that distribution, one word at a time.

We now know what we want to estimate. What remains is knowing how much uncertainty a distribution contains and how to measure how far a model is from reality. Claude Shannon answered those questions in 1948.

References

  1. J. Bernoulli (1713). Ars Conjectandi. Basel.
  2. T. Bayes (1763). “An Essay towards Solving a Problem in the Doctrine of Chances”. Philosophical Transactions of the Royal Society, 53.
  3. P.-S. Laplace (1814). Essai philosophique sur les probabilités. Paris.
  4. R. A. Fisher (1922). “On the Mathematical Foundations of Theoretical Statistics”. Philosophical Transactions of the Royal Society A, 222.
  5. E. T. Jaynes (2003). Probability Theory: The Logic of Science. Cambridge University Press.