Momentum and Adam

Level AdvancedDifficulty ★★★★★Application⌖ Open in the map

What is it?

Momentum averages past gradients (a heavy ball that keeps rolling through ravines); RMSProp divides by a running RMS of the gradient (per-parameter step sizes). Adam combines both and is the default optimizer of deep learning.

Formulas

mt=β1mt−1+(1−β1)gt,vt=β2vt−1+(1−β2)gt2,θt=θt−1−η m^tv^t+ϵm_t = \beta_1 m_{t-1} + (1-\beta_1)g_t, \quad v_t = \beta_2 v_{t-1} + (1-\beta_2)g_t^2, \quad \theta_t = \theta_{t-1} - \eta\,\frac{\hat m_t}{\sqrt{\hat v_t} + \epsilon}
Adam (with bias corrections m^,v^\hat m, \hat v)

The mathematics behind it

  • Variance★★★★★frequent

    Adam divides by a running estimate of the gradient's second moment.

  • Gradient descent with momentum is a discretized damped oscillator (heavy-ball ODE) rolling in the loss landscape.

Where is it used?

Computing topics reachable from here, through the chain of ideas that leads to them:

What depends on it

This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.

↑ ↓ to navigate · ↵ · Esc