Chapter 06 · Calculus
Which way does the error move?
A neural network has billions of knobs. To know which one to turn, and which way, you need to know how the error changes when each one moves: its derivative. The chain rule, which Leibniz wrote down in 1676, lets us compute them all at once. That algorithm is called backpropagation.
In this chapter
Newton and Leibniz invented calculus independently in the second half of the seventeenth century, to describe the motion of the planets and the shape of curves. Three centuries later, their central tool, the derivative, is what allows a neural network to learn. The idea is simple: if you know how the error changes when you nudge each parameter a little, you know how to move them so as to be less wrong.
The derivative: instantaneous sensitivity
The derivative of at a point measures how much changes when moves by a very small amount:
Geometrically, it is the slope of the tangent line. The quotient with a finite is the slope of a secant line, which crosses the curve at two points. As shrinks, the secant turns until it coincides with the tangent.
The gradient: the derivative with many variables
A loss function does not depend on one number but on millions: . Each partial derivative measures the sensitivity to one parameter with the others held fixed, and the vector collecting them all is the gradient:
If is differentiable at and , then among all unit directions the directional derivative is largest when points along and smallest when it points the opposite way.
Proof
By Cauchy–Schwarz, for , and the extremes are reached at .
Here is the recipe for learning: move in the direction of , the direction in which the error falls fastest. What we need is an efficient way to compute the gradient of a function with billions of variables.
The chain rule
A neural network is a composition of many simple functions, like an assembly line. The chain rule says how sensitivity propagates along that line. Leibniz used it in notes from 1676, and his notation makes it look like simply cancelling fractions:
If and are differentiable, then
With several variables the derivatives become Jacobian matrices and the product becomes a matrix product: .
For a long chain, , the derivative is a product of many factors. The order in which they are multiplied does not change the result, but it changes the cost of the computation enormously.
Computational graphs and backpropagation
Any computation can be drawn as a directed acyclic graph: the nodes are elementary operations and the edges carry the intermediate results. This is where graph theory enters AI. For a neuron that predicts and makes an error , the graph is the one in the figure.
This algorithm, reverse-mode automatic differentiation, was described by Seppo Linnainmaa in his 1970 master's thesis. Paul Werbos proposed applying it to neural networks in his 1974 thesis. The paper by David Rumelhart, Geoffrey Hinton and Ronald Williams in Nature (1986) showed that networks trained this way learn useful internal representations, and it popularized the method under the name backpropagation.
If a function is computed with elementary operations, its full gradient, all partial derivatives, can be computed in reverse mode with operations, where is a small constant (around 3 to 5) that does not depend on .
This is the result that makes deep learning possible. Computing each derivative separately, by moving one parameter and measuring the change, would cost evaluations of the network: with parameters, impossible. Backpropagation delivers every derivative for roughly the price of two or three evaluations. Every current AI library (PyTorch, JAX, TensorFlow) is, at heart, an automatic differentiation engine.
When the gradient vanishes
The chain rule has a dark side. The sigmoid , the classic activation function, has derivative . In a network with sigmoid layers, the gradient reaching the first layers contains a product of such factors:
If the weights are moderate, the gradient shrinks exponentially with depth and the first layers do not learn. Sepp Hochreiter analysed this vanishing gradient problem in his 1991 thesis. The solutions came in several waves: LSTMs (1997), the ReLU activation, , whose derivative is 1 on the positive side, and residual connections (2015), , which open a highway along which the gradient travels without fading. Transformers use the last two.
We now know how to compute which way the error falls. What remains is to decide how far to go in that direction, and what guarantees there are of getting anywhere. That is optimization.
References
- S. Linnainmaa (1970). The representation of the cumulative rounding error of an algorithm as a Taylor expansion of the local rounding errors. Master's thesis, University of Helsinki.
- P. J. Werbos (1974). Beyond Regression: New Tools for Prediction and Analysis in the Behavioral Sciences. PhD thesis, Harvard.
- W. Baur and V. Strassen (1983). “The complexity of partial derivatives”. Theoretical Computer Science, 22(3).
- D. E. Rumelhart, G. E. Hinton and R. J. Williams (1986). “Learning representations by back-propagating errors”. Nature, 323.
- S. Hochreiter (1991). Untersuchungen zu dynamischen neuronalen Netzen. Diploma thesis, TU Munich.
- A. Griewank and A. Walther (2008). Evaluating Derivatives, 2nd ed. SIAM.
- K. He, X. Zhang, S. Ren and J. Sun (2016). “Deep Residual Learning for Image Recognition”. CVPR.