What is it?
The geometry of over parameter space: valleys, plateaus, saddle points and ravines. Its curvature (Hessian) controls how fast and how stably optimizers move, and flat minima tend to generalize better than sharp ones.
Formulas
- sharpness
- the edge of stability, where full-batch training tends to settle
Why does it matter?
Understanding why non-convex training works so well is one of the open questions of deep learning theory.
The mathematics behind it
Hessian eigenvalues measure sharpness, detect saddles and set the stable learning rate .
Minima, maxima and (overwhelmingly) saddle points shape the landscape optimizers must cross.
Deep-network losses are non-convex: many minima and saddles, yet good minima are easy to find in practice.
A Hessian with large condition number makes gradient descent zigzag; preconditioning (Adam, normalization layers) helps.
Contour plots of a loss along two directions are the standard way to visualize its landscape.
In high dimension most critical points of a random-looking loss are saddles, not minima.
Where is it used?
Computing topics reachable from here, through the chain of ideas that leads to them:
What depends on it
This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.