What is it?
: the average squared distance from the mean. The variance of an average of independent samples is β the law that governs Monte Carlo error, batch sizes and A/B tests.
Formulas
Where it shows up in computing
MC error is ; variance reduction (importance sampling, control variates) is the main lever.
The Kalman gain weighs prediction and measurement by their variances.
Where it shows up in AI
Gradient noise variance falls as ; it sets the useful learning rate and batch size.
Adam divides by a running estimate of the gradient's second moment.
Xavier/He initialization chooses weight variances so activations keep a stable variance through layers.
Where is it used?
Computing topics reachable from here, through the chain of ideas that leads to them:
3D Computer graphics
- Monte Carlo methodsβThe rendering equationβ β β β β
β Robotics and control
- Kalman filterβ β β β β
β AI and machine learning
- Stochastic gradient descent (SGD)β β β β β
- Momentum and Adamβ β β β β
- Momentum and AdamβDeep learningβ β β β β
- Momentum and AdamβDeep learningβConvolutional networks (CNNs)β β β β β
- Momentum and AdamβDeep learningβGenerative modelsβ β β β β
- Momentum and AdamβDeep learningβNeural ODEsβ β β β β
- +4
This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.