What is it?
The step size of gradient descent — the most important hyperparameter. Stability requires (the largest curvature); schedules (warm-up, cosine decay) start careful, go fast, and end small to settle into a minimum.
Formulas
- cosine schedule
Where is it used?
Computing topics reachable from here, through the chain of ideas that leads to them:
ℒ AI and machine learning
- Stochastic gradient descent (SGD)★★★★★
- Stochastic gradient descent (SGD)→Momentum and Adam★★★★★
- Stochastic gradient descent (SGD)→Momentum and Adam→Deep learning★★★★★
- Stochastic gradient descent (SGD)→Momentum and Adam→Deep learning→Convolutional networks (CNNs)★★★★★
- Stochastic gradient descent (SGD)→Momentum and Adam→Deep learning→Generative models★★★★★
- Stochastic gradient descent (SGD)→Momentum and Adam→Deep learning→Neural ODEs★★★★★
What depends on it
This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.