What is it?
When a variable influences the output through several paths, add the contributions of every path, each the product of the local derivatives along it: . In matrix form, Jacobians multiply. This is exactly what backpropagation computes on a network's graph.
Why does it exist?
Real computations branch and merge: a weight affects many neurons, which all affect the loss. The one-variable chain rule handles a single path; the multivariable version handles any directed acyclic graph of operations — any program without loops, or any unrolled one.
Intuition
Draw the computation as a graph: inputs at the bottom, the output at the top, each edge labelled with a local partial derivative. is the sum, over all paths, of the products of labels along the path. Backpropagation computes all these sums at once by sweeping the graph from the top, accumulating at each node "how much the output cares about me" (the adjoint).
Formal definition
If is differentiable at and at , then
For scalar with : .
Formulas
- a deep network: a product of Jacobians
- the backpropagation equations
How is it computed?
The product can be evaluated right-to-left (forward mode: cost ∝ number of inputs) or left-to-right (reverse mode: cost ∝ number of outputs). A loss has one output and millions of inputs, so reverse mode wins by a factor of millions — that choice is backpropagation.
Example
with , . Two paths from to (through and through ): .
Why does it matter?
It is the mathematical content of backpropagation and of every autodiff library (PyTorch, JAX, TensorFlow). It also gives robot velocities through the kinematic chain and the sensitivities used in engineering design (adjoint methods in CFD and weather forecasting are the same reverse-mode idea).
Where it shows up in computing
4D-Var data assimilation uses adjoint (reverse-mode) models to fit the initial state of forecasts.
Where it shows up in AI
Backprop = the multivariable chain rule evaluated in reverse order on the network's graph.
Forward and reverse mode are the two natural orders of multiplying the Jacobian chain.
Where is it used?
Computing topics reachable from here, through the chain of ideas that leads to them:
λ Scientific computing and algorithms
- Jacobian matrix→Multiple integrals and change of variables→Monte Carlo methods★★★★★
- Jacobian matrix→Equilibria and stability→Stiffness and implicit methods→Scientific computing★★★★★
- Jacobian matrix→Equilibria and stability→Attractors→Chaos and sensitivity to initial conditions→Floating point (IEEE 754)★★★★★
⚛ Physics and simulation
- Jacobian matrix→Multiple integrals and change of variables→Surface integrals and flux→Divergence theorem (Gauss)→Electromagnetism (Maxwell's equations)★★★★★
- Jacobian matrix→Multiple integrals and change of variables→Surface integrals and flux→Divergence theorem (Gauss)→Fluid dynamics and CFD★★★★★
- Jacobian matrix→Equilibria and stability→Attractors→Chaos and sensitivity to initial conditions→Weather and climate modelling★★★★★
- Jacobian matrix→Equilibria and stability→Stiffness and implicit methods→Physics engines★★★★★
- Jacobian matrix→Equilibria and stability→Stiffness and implicit methods→Physics engines→N-body gravitational simulation★★★★★
- Jacobian matrix→Equilibria and stability→Bifurcations→Population and epidemic models★★★★★
What depends on it
Exercises
If with , , express and .
Solution
, .
A network computes . Write using the chain rule and say which factors are reused from .
Solution
Let , , . With : and . The upstream gradient is computed once and reused: that sharing is what makes backprop cheap.