Introduction
Why does AI work?
There is no single theory behind ChatGPT, but a convergence of many: logic, probability, information, linear algebra, calculus, optimization… This site walks through them in order, from Boole and Turing to Transformers, with the theorems that matter, the history of the people who discovered them, and an interactive figure at every step.
Start at the beginning →See the timeline
The idea in one sentence
A language model learns to predict which information is most likely to come next. It needs no explicit definition of a cat to complete “The cat is sitting on the…”. It is enough to have estimated, from enormous amounts of text, an extraordinarily detailed probability distribution. That something so simple to state gives such astonishing results has precise mathematical explanations (and a few questions that are still open). Here is the journey:
-
00
Switch 1854 – 1959
How does a machine compute, deep down?
Underneath every program there are switches. Boolean algebra, a few logic gates and a little memory are enough to build an adder, a finite automaton and finally a processor. Along the way the first limit appears: there are patterns that no machine with finite memory can recognise.
- Functional completeness of NAND
- Kleene's theorem
- Subset construction
- Pumping lemma
- Chomsky hierarchy
-
01
Compute 1928 – 1950
What can a machine compute?
Before asking whether a machine can think, someone had to define what it means to compute. Alan Turing did it with a tape, a head and a table of rules, and along the way discovered that there are questions no computer will ever be able to answer.
- Universal machine
- Undecidability of halting
- Church–Turing thesis
-
02
Deduce 1854 – 1992
Can a machine reason?
The first plan for artificial intelligence was not to learn but to deduce: write knowledge down as logical formulas and let the machine draw the consequences. It produced a beautiful theory, a programming language in which you describe the problem instead of the solution, and a lesson about its limits.
- Gödel's completeness
- Most general unifier
- Completeness of resolution
- Least Herbrand model
- Completeness of SLD resolution
-
03
Probability 1654 – 1922
Learning from data
Probability puts numbers on uncertainty. Two eighteenth-century results, the law of large numbers and Bayes' theorem, explain why more data brings us closer to the truth and how to update what we believe as each new data point arrives.
- Law of large numbers
- Bayes' theorem
- Maximum likelihood
- Chain rule of probability
-
04
Information 1948 – 1952
Measuring surprise
In 1948 Claude Shannon turned information into a measurable quantity. With a single formula, entropy, he explained how far a message can be compressed and how to measure how far a model is from reality. That measure is, literally, the function every LLM minimizes during training.
- Shannon entropy
- Source coding theorem
- Gibbs' inequality
- Cross-entropy
-
05
Vectors 1844 – 2013
Words turned into arrows
A computer has no idea what a cat is, but it can multiply matrices at enormous speed. Linear algebra lets us represent words, images and concepts as vectors, and measure how alike they are with a simple dot product.
- Cauchy–Schwarz inequality
- Johnson–Lindenstrauss lemma
- Singular value decomposition
-
06
Derivatives 1676 – 1986
Which way does the error move?
A neural network has billions of knobs. To know which one to turn, and which way, you need to know how the error changes when each one moves: its derivative. The chain rule, which Leibniz wrote down in 1676, lets us compute them all at once. That algorithm is called backpropagation.
- Chain rule
- Direction of steepest ascent
- Cheap gradient principle
-
07
Optimize 1847 – 2014
Walking down a mountain blindfolded
Training a model means searching for the lowest point of a landscape with billions of dimensions, in thick fog, feeling only the slope underfoot. Gradient descent, invented by Cauchy in 1847 to compute orbits, is still the algorithm that trains every LLM.
- Convergence of gradient descent
- Minima of convex functions
- Robbins–Monro conditions
- Nesterov's lower bound
-
08
Neurons 1943 – 2012
From the perceptron to universal approximation
In 1958 a US Navy machine learned to tell cards marked on the left from cards marked on the right, and the press announced it would soon walk, talk and be conscious of its existence. Eleven years later, a book proved it could not compute something as simple as “exclusive or”. This is the story of how a single hidden layer fixed everything.
- Perceptron convergence
- XOR is not linearly separable
- Universal approximation
-
09
Generalize 1971 – 2019
Why does it work on data it has never seen?
Memorizing the examples is easy; getting unseen ones right is not. Learning theory studies when and why a model trained on some data generalizes to other data. Its classical theorems are elegant… and deep networks defy them.
- Generalization bound (finite class)
- VC dimension
- Fundamental theorem of PAC learning
- No free lunch
-
10
Attention 1990 – 2017
Attention: every word looks at all the others
In 2017 eight researchers at Google published a paper with a provocative title: “Attention Is All You Need”. Their architecture, the Transformer, replaced word-by-word reading with an operation in which each word decides, with a dot product and a softmax, which others to pay attention to.
- Attention as soft lookup
- Variance of the dot product ()
- Universal approximation of sequences
-
11
LLM 2018 – today
Predicting the next word, at planetary scale
An LLM is a Transformer trained to minimize Shannon's cross-entropy over trillions of words. This chapter brings every piece together: how it generates text, why it improves predictably as it grows, how it becomes an assistant like ChatGPT, and why it sometimes makes things up.
- Softmax as maximum entropy
- Scaling laws
- Optimal policy with KL
- Calibration and hallucinations
Many theories, one chain
Each chapter answers one question and raises the next. Several branches of mathematics and computer science appear along the way:
| Theory | What it brings to modern AI | Where |
|---|---|---|
| Boolean algebra and automata | Circuits, finite memory and what it cannot recognise | 00 |
| Theory of computation | What a machine can compute, and at what cost | 01 |
| Logic | Deduction, resolution and logic programming | 02 |
| Probability and statistics | Prediction, uncertainty and learning from data | 03 |
| Information theory | Entropy, compression and the loss function of LLMs | 04 |
| Linear algebra | Vectors, matrices and embeddings: the space where meaning lives | 05 |
| Calculus | Derivatives, the chain rule and backpropagation | 06 |
| Graph theory | Computational graphs along which the gradient flows | 06 |
| Optimization | Training by gradient descent | 07 |
| Neural networks | Functions able to approximate almost anything | 08 |
| Learning theory | Why and when a machine generalizes | 09 |
| Information and attention | The central mechanism of the Transformer | 10 |
| Games and reinforcement learning | Agents that learn from rewards: RLHF | 11 |
How to read this site
You can read it in order, like a short book, or jump to the chapter you are interested in: each one stands on its own and links back to earlier ones when needed. Proofs are folded so they don't interrupt the reading; open them if you want to see why something is true. The figures are interactive and run in your browser, with no servers and no trackers. High-school science is enough to follow the thread, and anyone who wants more detail will find the original references in every chapter.