What is it?
Momentum averages past gradients (a heavy ball that keeps rolling through ravines); RMSProp divides by a running RMS of the gradient (per-parameter step sizes). Adam combines both and is the default optimizer of deep learning.
Formulas
- Adam (with bias corrections )
The mathematics behind it
Adam divides by a running estimate of the gradient's second moment.
Gradient descent with momentum is a discretized damped oscillator (heavy-ball ODE) rolling in the loss landscape.
Where is it used?
Computing topics reachable from here, through the chain of ideas that leads to them:
What depends on it
This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.