Loss landscape

Level AdvancedDifficulty ★★★★★Application⌖ Open in the map

What is it?

The geometry of L(θ)L(\theta) over parameter space: valleys, plateaus, saddle points and ravines. Its curvature (Hessian) controls how fast and how stably optimizers move, and flat minima tend to generalize better than sharp ones.

Formulas

S(θ)=λmax⁡(∇2L(θ))S(\theta) = \lambda_{\max}\big(\nabla^2 L(\theta)\big)
sharpness
η λmax⁡≈2\eta\,\lambda_{\max} \approx 2
the edge of stability, where full-batch training tends to settle

Why does it matter?

Understanding why non-convex training works so well is one of the open questions of deep learning theory.

The mathematics behind it

  • Hessian matrix★★★★★fundamental

    Hessian eigenvalues measure sharpness, detect saddles and set the stable learning rate η<2/λmax⁡\eta < 2/\lambda_{\max}.

  • Critical points and derivative tests★★★★★fundamental

    Minima, maxima and (overwhelmingly) saddle points shape the landscape optimizers must cross.

  • Convexity and concavity★★★★★frequent

    Deep-network losses are non-convex: many minima and saddles, yet good minima are easy to find in practice.

  • Conditioning★★★★★frequent

    A Hessian with large condition number makes gradient descent zigzag; preconditioning (Adam, normalization layers) helps.

  • Surfaces and level sets★★★★★frequent

    Contour plots of a loss along two directions are the standard way to visualize its landscape.

  • Extrema in several variables★★★★★frequent

    In high dimension most critical points of a random-looking loss are saddles, not minima.

Where is it used?

Computing topics reachable from here, through the chain of ideas that leads to them:

What depends on it

This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.

↑ ↓ to navigate · ↵ · Esc