What is it?
: the vector of all partial derivatives. It points in the direction of steepest ascent, its length is that steepest slope, and it is perpendicular to the level sets. Walk against it and you go downhill fastest.
Why does it exist?
We want one object that answers "which way is up, and how steep?" for a function of many variables. The partial derivatives are separate numbers; assembled into a vector they gain a geometric meaning that does not depend on the coordinate system — and that turns optimization into "follow an arrow".
Intuition
Stand on a hillside in fog. The gradient is the arrow on the ground pointing straight uphill; contour lines run perpendicular to it. Its length is the slope in that direction. Gradient descent is: look at the arrow, take a step the other way, repeat.
∇f
↑
│
╭───────●───────╮ ← level curve f = c
╭─┴───────────────┴─╮
Formal definition
For differentiable at , is the unique vector with
The directional derivative in a unit direction is , maximized by (Cauchy–Schwarz), with maximum value . If , it is orthogonal to the level set .
Formulas
- steepest ascent along
- gradient descent
- unit normal of an implicit surface
How is it computed?
By hand: all partial derivatives. In software, for a scalar loss of parameters, reverse-mode automatic differentiation computes the whole gradient for a small constant times the cost of evaluating (Baur–Strassen), while finite differences would need evaluations.
Example
(an elongated bowl). . At : , which does not point at the minimum — it points across the narrow valley. That is why plain gradient descent zigzags on badly conditioned problems (try it in the Gradient descent demo).
Why does it matter?
The gradient is the bridge from calculus to modern AI: every model trained by gradient descent — from linear regression to large language models — follows . It is also how renderers get surface normals from implicit shapes, how edge detectors find boundaries, and how physics derives forces from potentials ().
Where it shows up in computing
The normal of an implicit surface is .
For a signed distance , and is the surface normal; ray marchers estimate it by finite differences.
Edges are large ; HOG features are histograms of gradient directions.
Conservative forces are minus the gradient of a potential: .
Where it shows up in AI
The update is the gradient, used as a direction.
Where is it used?
Computing topics reachable from here, through the chain of ideas that leads to them:
What depends on it
Exercises
Find for at and the rate of increase in the direction .
Solution
. ; the maximum possible is .
Draw the level curves of and the gradient at . Why does the arrow not point to the origin?
Solution
Level curves are ellipses, wider in . The gradient is perpendicular to the ellipse through , which is not the radial direction because the ellipse is not a circle.
Why does reverse-mode AD compute in roughly the time of 3–5 evaluations of , while finite differences would need ?
Solution
Finite differences perturb one parameter at a time. Reverse mode runs the program once forward, then once backward propagating ; each intermediate is visited once, so the cost is a constant multiple of the forward pass, independent of the number of parameters (one scalar output).