What is it?
The derivative of a composition is the product of the derivatives: . Rates of change multiply along a chain — and backpropagation is this rule applied, very efficiently, to a neural network.
Why does it exist?
Almost every function we meet is built by composing simpler ones. Without the chain rule we would have to go back to limits for each new formula; with it, knowing the derivatives of a few primitives is enough to differentiate everything built from them — including a program.
Intuition
Gears. If turns 3 times as fast as , and turns 2 times as fast as , then turns times as fast as . In Leibniz notation the rule looks like cancelling fractions: . Along a chain of functions you multiply local rates — and if they are all smaller than 1, the product vanishes (that is the vanishing-gradient problem).
Formal definition
If is differentiable at and is differentiable at , then is differentiable at and
Idea of the proof
Write with and ; divide by and let . (Dividing and multiplying by directly fails when , which is why the auxiliary function is needed.)
Formulas
- a chain of functions: a product of local derivatives
How is it computed?
Name the intermediate values: , , … Compute the local derivatives and multiply. The order in which you multiply does not change the result, but it changes the cost when the functions have many inputs and outputs: left-to-right is forward mode, right-to-left is reverse mode (backpropagation).
Example
with the sigmoid. Set . Then . This is literally the gradient a neural network computes for one weight of one neuron.
Why does it matter?
Backpropagation (Rumelhart, Hinton and Williams, 1986; Linnainmaa's reverse-mode AD, 1970) is the chain rule organised so that the gradient with respect to all weights costs about as much as one evaluation of the network. Without that trick, training modern models would be impossible.
Where it shows up in computing
End-effector velocity is the chain rule through the arm: .
Where it shows up in AI
Backprop is the chain rule evaluated from the loss backwards, reusing every intermediate product.
Forward and reverse mode are two orders of multiplying the same chain of local derivatives.
Where is it used?
Computing topics reachable from here, through the chain of ideas that leads to them:
⚙ Robotics and control
- Multivariable chain rule→Jacobian matrix→Robot Jacobian (velocity kinematics)★★★★★
- Multivariable chain rule→Jacobian matrix→Inverse kinematics★★★★★
- Multivariable chain rule→Jacobian matrix→Equilibria and stability→Control theory★★★★★
- Multivariable chain rule→Jacobian matrix→Equilibria and stability→Control theory→Trajectory optimization and MPC★★★★★
- Multivariable chain rule→Jacobian matrix→Kalman filter★★★★★
⚛ Physics and simulation
- Integration by substitution→Trigonometric integrals→Fourier series→Heat equation and diffusion★★★★★
- Integration by substitution→Separable equations→Population and epidemic models★★★★★
- Multivariable chain rule→Jacobian matrix→Equilibria and stability→Stiffness and implicit methods→Physics engines★★★★★
- Multivariable chain rule→Jacobian matrix→Multiple integrals and change of variables→Surface integrals and flux→Fluid dynamics and CFD★★★★★
- Multivariable chain rule→Jacobian matrix→Equilibria and stability→Attractors→Weather and climate modelling★★★★★
∿ Signals, media and vision
- Integration by substitution→Trigonometric integrals→Fourier series→Signal processing★★★★★
- Integration by substitution→Trigonometric integrals→Fourier series→Fourier transform→Sampling theorem (Nyquist–Shannon)★★★★★
- Integration by substitution→Trigonometric integrals→Fourier series→Fourier transform→Fast Fourier transform (FFT)★★★★★
- Integration by substitution→Trigonometric integrals→Fourier series→Signal processing→Digital filters★★★★★
- Integration by substitution→Trigonometric integrals→Fourier series→Signal processing→Media compression (JPEG, MP3, video)★★★★★
- Integration by substitution→Trigonometric integrals→Fourier series→Fourier transform→Telecommunications (modulation, OFDM)★★★★★
- +1
What depends on it
Exercises
Differentiate .
Solution
— the derivative of softplus is a sigmoid.
A 20-layer network uses sigmoid activations. Bound the factor by which the gradient can shrink when it goes through the 20 activations, ignoring the weights.
Solution
Each , so the product is at most : vanishing gradients. ReLU ( when active) and residual connections avoid it.