Upstream
Where does backpropagation come from?
- Functions
- Derivative
- Chain rule
- Functions of several variables
- Partial derivatives
- Gradient
- Multivariable chain rule
- Loss function
- Gradient descent
- Neural networks
- Backpropagation
Overview
All of university calculus — from limits to vector calculus, differential equations and Fourier — and where every idea shows up in AI, graphics, simulation, robotics and scientific computing.
Upstream
Downstream
The classic first-year sequence, from functions to several variables, for anyone who wants the foundations back.
∇The shortest honest path from derivatives to deep learning: gradients, optimization, likelihood and backpropagation.
3DFrom derivatives to rendering: vectors, normals, surfaces, lighting and the integral that gives each pixel its colour.
ẏHow a computer predicts motion: differential equations, numerical integrators, stability and chaos.
⚛Fields and waves: vector calculus, Fourier analysis and partial differential equations, up to fluids.
is the instantaneous rate of change of at : the slope of the tangent line to the graph, defined as the limit of slopes of secant lines.
Approximate the area under a curve by thin rectangles, . As the sum converges to the integral — slowly for the left/right rule (), faster for the midpoint ().
To solve , replace by its tangent line at the current guess and jump to where the tangent hits zero: . Near a simple root the number of correct digits doubles every step.
The polynomial of degree that matches and its first derivatives at a point . Degree 1 is the tangent line, degree 2 adds curvature; the higher the degree, the wider the region where it is a good approximation. At it is called a Maclaurin polynomial.
Repeat : take a small step against the gradient. Cauchy proposed it in 1847; today it (and its stochastic, adaptive variants) trains essentially every neural network.
The algorithm that computes and for every layer of a network: one forward pass storing intermediate values, then one backward pass applying the chain rule from the loss down to the inputs. Cost: about twice the forward pass, whatever the number of parameters.