Chapter 03 · Probability
Learning from data
Probability puts numbers on uncertainty. Two eighteenth-century results, the law of large numbers and Bayes' theorem, explain why more data brings us closer to the truth and how to update what we believe as each new data point arrives.
In this chapter
In 1654 the Chevalier de Méré put a gambling problem to Blaise Pascal: if a game of several rounds is interrupted before it ends, how should the stakes be fairly divided? Pascal wrote about it to Pierre de Fermat, and the calculus of probability was born in that correspondence. Three centuries later, the same mathematics that split wagers decides which word ChatGPT writes next.
A language model does not “know” which word comes next: it assigns a probability to every possible word. Everything it does rests on this language.
Random variables and distributions
A random variable is a quantity whose value depends on chance: the outcome of a coin, the height of a randomly chosen person, or the next word of a text. Its distribution assigns each outcome a number between 0 and 1, and they all add up to 1:
The expected value is the average weighted by the probabilities, , and the variance measures how much the variable spreads around that mean, .
The law of large numbers
If a coin lands heads with a probability we do not know, intuition says that after many flips the proportion of heads will approach . Jakob Bernoulli spent twenty years proving it, and the result was published in 1713, eight years after his death, in Ars Conjectandi. He called it his “golden theorem”.
Let be independent variables with the same distribution, mean and finite variance , and let be their sample mean. For every ,
Proof (with Chebyshev's inequality)
By independence, . Chebyshev's inequality says that a variable strays from its mean by more than with probability at most , so
The proof gives more than the limit: the typical error of the mean shrinks like . To be twice as precise you need four times as much data. This law of diminishing returns reappears, with other exponents, in the scaling laws of LLMs.
Bayes' theorem
The law of large numbers speaks about what happens in the long run, but we almost never have infinite data. What can we say about after only ten flips? The Reverend Thomas Bayes tackled this inverse problem in an essay his friend Richard Price published in 1763, after his death. Laplace rediscovered and generalized it in 1774.
For a hypothesis and data with ,
Proof
By the definition of conditional probability, . Just solve for .
The theorem fits on one line, but its reading is deep: the prior summarizes what we believed before seeing the data, the likelihood measures how well each hypothesis explains what was observed, and the posterior is the updated belief. Each new data point turns today's posterior into tomorrow's prior.
For the coin, if the prior over is a distribution and we observe heads in flips, the posterior is again a Beta:
With the uniform prior (), the probability that the next flip lands heads is . This is Laplace's rule of succession, with which in 1814 he computed the probability that the sun will rise tomorrow given that it has risen every day of recorded history. It was the first “smoothing” of an estimate, and statistical language models would use it two hundred years later so as not to assign zero probability to words they had never seen.
Maximum likelihood: the loss function of AI
Between 1912 and 1922 Ronald Fisher proposed a more direct criterion: among all possible values of the parameter, keep the one that makes the observed data most probable. If the data are independent the likelihood is a product, and since multiplying many small numbers causes numerical trouble, its logarithm is maximized instead:
For the coin we get the expected . But the expression on the right deserves a second look: it is exactly the function minimized when training any language model. In the next chapter we will see that it also has a name in information theory: cross-entropy.
Training a model means choosing the parameters that make the training data most probable. Everything else (networks, gradients, attention) is a way of expressing and optimizing that probability.
The chain rule of probability
One piece is missing. How do you assign a probability to a whole sentence, when the number of possible sentences is astronomical? With an elementary identity that follows from applying the definition of conditional probability over and over:
The probability of a text is the product of the probabilities of each word conditioned on all the previous ones. This identity turns the impossible problem of modelling whole texts into a manageable one: predicting the next word. An LLM is a gigantic approximator of , and it generates text by sampling from that distribution, one word at a time.
We now know what we want to estimate. What remains is knowing how much uncertainty a distribution contains and how to measure how far a model is from reality. Claude Shannon answered those questions in 1948.
References
- J. Bernoulli (1713). Ars Conjectandi. Basel.
- T. Bayes (1763). “An Essay towards Solving a Problem in the Doctrine of Chances”. Philosophical Transactions of the Royal Society, 53.
- P.-S. Laplace (1814). Essai philosophique sur les probabilités. Paris.
- R. A. Fisher (1922). “On the Mathematical Foundations of Theoretical Statistics”. Philosophical Transactions of the Royal Society A, 222.
- E. T. Jaynes (2003). Probability Theory: The Logic of Science. Cambridge University Press.