Introduction

Why does AI work?

There is no single theory behind ChatGPT, but a convergence of many: logic, probability, information, linear algebra, calculus, optimization… This site walks through them in order, from Boole and Turing to Transformers, with the theorems that matter, the history of the people who discovered them, and an interactive figure at every step.

Start at the beginning →See the timeline

This is how a language model “writes”: at every step it computes a probability distribution over the next word and picks one. Choose words yourself or let chance do it, and watch three numbers: the entropy (how much uncertainty there is before choosing), the surprise of what was chosen, and the probability of the sentence, which is the product of the probabilities of each step. The rest of the site explains where those numbers come from and why they work.

The idea in one sentence

A language model learns to predict which information is most likely to come next. It needs no explicit definition of a cat to complete “The cat is sitting on the…”. It is enough to have estimated, from enormous amounts of text, an extraordinarily detailed probability distribution. That something so simple to state gives such astonishing results has precise mathematical explanations (and a few questions that are still open). Here is the journey:

  1. 00 Switch 1854 – 1959 How does a machine compute, deep down? Underneath every program there are switches. Boolean algebra, a few logic gates and a little memory are enough to build an adder, a finite automaton and finally a processor. Along the way the first limit appears: there are patterns that no machine with finite memory can recognise.
    • Functional completeness of NAND
    • Kleene's theorem
    • Subset construction
    • Pumping lemma
    • Chomsky hierarchy
  2. 01 Compute 1928 – 1950 What can a machine compute? Before asking whether a machine can think, someone had to define what it means to compute. Alan Turing did it with a tape, a head and a table of rules, and along the way discovered that there are questions no computer will ever be able to answer.
    • Universal machine
    • Undecidability of halting
    • Church–Turing thesis
  3. 02 Deduce 1854 – 1992 Can a machine reason? The first plan for artificial intelligence was not to learn but to deduce: write knowledge down as logical formulas and let the machine draw the consequences. It produced a beautiful theory, a programming language in which you describe the problem instead of the solution, and a lesson about its limits.
    • Gödel's completeness
    • Most general unifier
    • Completeness of resolution
    • Least Herbrand model
    • Completeness of SLD resolution
  4. 03 Probability 1654 – 1922 Learning from data Probability puts numbers on uncertainty. Two eighteenth-century results, the law of large numbers and Bayes' theorem, explain why more data brings us closer to the truth and how to update what we believe as each new data point arrives.
    • Law of large numbers
    • Bayes' theorem
    • Maximum likelihood
    • Chain rule of probability
  5. 04 Information 1948 – 1952 Measuring surprise In 1948 Claude Shannon turned information into a measurable quantity. With a single formula, entropy, he explained how far a message can be compressed and how to measure how far a model is from reality. That measure is, literally, the function every LLM minimizes during training.
    • Shannon entropy
    • Source coding theorem
    • Gibbs' inequality
    • Cross-entropy
  6. 05 Vectors 1844 – 2013 Words turned into arrows A computer has no idea what a cat is, but it can multiply matrices at enormous speed. Linear algebra lets us represent words, images and concepts as vectors, and measure how alike they are with a simple dot product.
    • Cauchy–Schwarz inequality
    • Johnson–Lindenstrauss lemma
    • Singular value decomposition
  7. 06 Derivatives 1676 – 1986 Which way does the error move? A neural network has billions of knobs. To know which one to turn, and which way, you need to know how the error changes when each one moves: its derivative. The chain rule, which Leibniz wrote down in 1676, lets us compute them all at once. That algorithm is called backpropagation.
    • Chain rule
    • Direction of steepest ascent
    • Cheap gradient principle
  8. 07 Optimize 1847 – 2014 Walking down a mountain blindfolded Training a model means searching for the lowest point of a landscape with billions of dimensions, in thick fog, feeling only the slope underfoot. Gradient descent, invented by Cauchy in 1847 to compute orbits, is still the algorithm that trains every LLM.
    • Convergence of gradient descent
    • Minima of convex functions
    • Robbins–Monro conditions
    • Nesterov's lower bound
  9. 08 Neurons 1943 – 2012 From the perceptron to universal approximation In 1958 a US Navy machine learned to tell cards marked on the left from cards marked on the right, and the press announced it would soon walk, talk and be conscious of its existence. Eleven years later, a book proved it could not compute something as simple as “exclusive or”. This is the story of how a single hidden layer fixed everything.
    • Perceptron convergence
    • XOR is not linearly separable
    • Universal approximation
  10. 09 Generalize 1971 – 2019 Why does it work on data it has never seen? Memorizing the examples is easy; getting unseen ones right is not. Learning theory studies when and why a model trained on some data generalizes to other data. Its classical theorems are elegant… and deep networks defy them.
    • Generalization bound (finite class)
    • VC dimension
    • Fundamental theorem of PAC learning
    • No free lunch
  11. 10 Attention 1990 – 2017 Attention: every word looks at all the others In 2017 eight researchers at Google published a paper with a provocative title: “Attention Is All You Need”. Their architecture, the Transformer, replaced word-by-word reading with an operation in which each word decides, with a dot product and a softmax, which others to pay attention to.
    • Attention as soft lookup
    • Variance of the dot product (dk)
    • Universal approximation of sequences
  12. 11 LLM 2018 – today Predicting the next word, at planetary scale An LLM is a Transformer trained to minimize Shannon's cross-entropy over trillions of words. This chapter brings every piece together: how it generates text, why it improves predictably as it grows, how it becomes an assistant like ChatGPT, and why it sometimes makes things up.
    • Softmax as maximum entropy
    • Scaling laws
    • Optimal policy with KL
    • Calibration and hallucinations

Many theories, one chain

Each chapter answers one question and raises the next. Several branches of mathematics and computer science appear along the way:

TheoryWhat it brings to modern AIWhere
Boolean algebra and automataCircuits, finite memory and what it cannot recognise00
Theory of computationWhat a machine can compute, and at what cost01
LogicDeduction, resolution and logic programming02
Probability and statisticsPrediction, uncertainty and learning from data03
Information theoryEntropy, compression and the loss function of LLMs04
Linear algebraVectors, matrices and embeddings: the space where meaning lives05
CalculusDerivatives, the chain rule and backpropagation06
Graph theoryComputational graphs along which the gradient flows06
OptimizationTraining by gradient descent07
Neural networksFunctions able to approximate almost anything08
Learning theoryWhy and when a machine generalizes09
Information and attentionThe central mechanism of the Transformer10
Games and reinforcement learningAgents that learn from rewards: RLHF11

How to read this site

You can read it in order, like a short book, or jump to the chapter you are interested in: each one stands on its own and links back to earlier ones when needed. Proofs are folded so they don't interrupt the reading; open them if you want to see why something is true. The figures are interactive and run in your browser, with no servers and no trackers. High-school science is enough to follow the thread, and anyone who wants more detail will find the original references in every chapter.

Chapter 00: How does a machine compute, deep down? →