What is it?
Computing exact derivatives of programs by applying the chain rule to each elementary operation. Neither symbolic (no expression swell) nor numerical (no truncation error). Reverse mode gives the gradient of a scalar for a small constant times the cost of the program.
Formulas
- forward mode with dual numbers
- reverse mode: adjoints propagate backwards
Why does it matter?
Every deep learning framework is, at its core, an AD engine (PyTorch autograd, JAX, TensorFlow). AD also powers differentiable rendering, differentiable physics and sensitivity analysis in science.
The mathematics behind it
AD computes exact derivatives of programs by propagating them through each operation.
AD applies exactly these rules, one elementary operation at a time, to numbers instead of formulas.
Forward and reverse mode are two orders of multiplying the same chain of local derivatives.
Forward and reverse mode are the two natural orders of multiplying the Jacobian chain.
AD sees a program as a composition of primitive operations and differentiates each one.
Every AD system ships a table of derivatives of its primitive operations.
Forward-mode AD with dual numbers computes (a JVP) in one pass.
AD computes JVPs (forward mode) and VJPs (reverse mode) without materializing .
"Gradient checking" compares autodiff gradients against finite differences to catch bugs.
Where is it used?
Computing topics reachable from here, through the chain of ideas that leads to them:
What depends on it
This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.