What is it?
Use the gradient of a small random batch instead of the whole dataset: a noisy but unbiased estimate, thousands of times cheaper. The noise even helps escape saddles and sharp minima.
Formulas
- Robbins–Monro step-size conditions for convergence
The mathematics behind it
A mini-batch gradient is an unbiased estimate of the expected gradient.
Gradient noise variance falls as ; it sets the useful learning rate and batch size.
Where is it used?
Computing topics reachable from here, through the chain of ideas that leads to them:
What depends on it
This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.