What is it?
The matrix of second partial derivatives : the curvature of in every direction. Its eigenvalues classify critical points (minimum, maximum, saddle) and control how fast optimizers can go.
Why does it exist?
The gradient says which way is downhill but not how the slope changes — whether the valley is a gentle bowl, a narrow ravine or a saddle. That second-order information decides whether a critical point is a minimum and how big a step is safe.
Intuition
Near a critical point, : a quadratic bowl whose axes are the eigenvectors of and whose steepness along each axis is the eigenvalue. All positive: a bowl (minimum). All negative: a dome. Mixed signs: a saddle. Very different eigenvalues: a narrow ravine where gradient descent bounces from wall to wall.
Formal definition
For , , symmetric by Schwarz. At a critical point : (positive definite) ⇒ strict local minimum; ⇒ strict local maximum; indefinite ⇒ saddle point.
Formulas
- conditioning and the largest stable learning rate on a quadratic
Example
: at the origin, , indefinite: a saddle (a Pringles chip). Gradient descent started exactly on the -axis converges to it; any tiny perturbation in escapes. That is why saddle points slow training down but rarely trap it.
Why does it matter?
Newton's method uses ; quasi-Newton methods (BFGS, L-BFGS) approximate it; Adam and other adaptive optimizers can be read as cheap diagonal approximations of curvature. In deep learning, the "sharpness" (largest Hessian eigenvalue) is linked to generalization and to the edge-of-stability phenomenon. The full Hessian of a billion-parameter model is never formed, but Hessian–vector products are cheap with autodiff.
Where it shows up in computing
Hessian-based detectors (determinant of Hessian in SURF, Frangi vesselness) find blobs and ridges.
Where it shows up in AI
Newton, Gauss–Newton, natural gradient and K-FAC use the Hessian or an approximation of it.
Hessian eigenvalues measure sharpness, detect saddles and set the stable learning rate .
Where is it used?
Computing topics reachable from here, through the chain of ideas that leads to them:
ℒ AI and machine learning
- Loss landscape★★★★★
- Second-order (Hessian-based) optimization★★★★★
- Extrema in several variables→Linear regression★★★★★
- Extrema in several variables→Linear regression→Logistic regression★★★★★
- Extrema in several variables→Linear regression→Logistic regression→Neural networks★★★★★
- Extrema in several variables→Linear regression→Logistic regression→Neural networks→Backpropagation★★★★★
What depends on it
Exercises
Classify the critical points of .
Solution
at . : at positive definite → minimum; at indefinite → saddle.
On , what is the largest learning rate for which gradient descent converges? How many steps to reduce the -error by at that rate?
Solution
, so . Along each step multiplies the error by : . Condition number 100 makes it slow.