∇ Mathematics for AI
The shortest honest path from derivatives to deep learning: gradients, optimization, likelihood and backpropagation.
- 01
Functions
A rule that assigns to each input of a set exactly one output in a set . Calculus studies how outputs change when inputs change; computing is, quite literally, evaluating functions.
- 02
Derivative
is the instantaneous rate of change of at : the slope of the tangent line to the graph, defined as the limit of slopes of secant lines.
- 03
Chain rule
The derivative of a composition is the product of the derivatives: . Rates of change multiply along a chain — and backpropagation is this rule applied, very efficiently, to a neural network.
- 04
Partial derivatives
: the derivative with respect to one variable, holding the others fixed. Each answers "how sensitive is the output to this input?" — for a neural network, to this weight.
- 05
Gradient
: the vector of all partial derivatives. It points in the direction of steepest ascent, its length is that steepest slope, and it is perpendicular to the level sets. Walk against it and you go downhill fastest.
- 06
Multivariable chain rule
When a variable influences the output through several paths, add the contributions of every path, each the product of the local derivatives along it: . In matrix form, Jacobians multiply. This is exactly what backpropagation computes on a network's graph.
- 07
Convexity and concavity
is convex if the chord between any two points of its graph lies above the graph — equivalently, for smooth , if . For convex functions every local minimum is global, which is why convex problems are the ones optimization can reliably solve.
- 08
Gradient descent
Repeat : take a small step against the gradient. Cauchy proposed it in 1847; today it (and its stochastic, adaptive variants) trains essentially every neural network.
- 09
Probability density function
A function with whose integral over a set is the probability of that set. A density is not a probability: it can exceed 1; it is probability per unit length.
- 10
Maximum likelihood estimation
Choose the parameters that make the observed data most probable: maximize , i.e. minimize . Squared error, cross-entropy and the training objective of language models are all negative log-likelihoods.
- 11
Loss function
A single number measuring how wrong a model with parameters is on the data. Learning is minimizing it. Mean squared error for regression and cross-entropy for classification are the two everyone uses.
- 12
Neural networks
Compositions of affine maps and non-linearities, with . A differentiable function with millions of adjustable parameters, trained by gradient descent.
- 13
Backpropagation
The algorithm that computes and for every layer of a network: one forward pass storing intermediate values, then one backward pass applying the chain rule from the loss down to the inputs. Cost: about twice the forward pass, whatever the number of parameters.
- 14
Deep learning
Training very deep networks (CNNs, transformers) with backpropagation and adaptive SGD on huge datasets. The calculus is the same as for one neuron; what changed is scale, architecture (residual connections, normalization, attention) and hardware.