Variance

Level UniversityDifficulty β˜…β˜…β˜…β˜…β˜…ConceptβŒ– Open in the map

What is it?

Var⁑(X)=𝔼[(Xβˆ’π”ΌX)2]\Var(X) = \E[(X - \E X)^2]: the average squared distance from the mean. The variance of an average of NN independent samples is Οƒ2/N\sigma^2/N β€” the 1/N1/\sqrt N law that governs Monte Carlo error, batch sizes and A/B tests.

Formulas

Var⁑(X)=𝔼[X2]βˆ’(𝔼X)2,Var⁑(1Nβˆ‘i=1NXi)=Οƒ2N\Var(X) = \E[X^2] - (\E X)^2, \qquad \Var\Big(\frac1N\sum_{i=1}^{N}X_i\Big) = \frac{\sigma^2}{N}

Where it shows up in computing

  • Monte Carlo methodsβ˜…β˜…β˜…β˜…β˜…fundamentalScientific computing and algorithms

    MC error is Οƒ/N\sigma/\sqrt N; variance reduction (importance sampling, control variates) is the main lever.

  • Kalman filterβ˜…β˜…β˜…β˜…β˜…fundamentalRobotics and control

    The Kalman gain weighs prediction and measurement by their variances.

Where it shows up in AI

  • Stochastic gradient descent (SGD)β˜…β˜…β˜…β˜…β˜…frequentAI and machine learning

    Gradient noise variance falls as 1/∣B∣1/|B|; it sets the useful learning rate and batch size.

  • Momentum and Adamβ˜…β˜…β˜…β˜…β˜…frequentAI and machine learning

    Adam divides by a running estimate of the gradient's second moment.

  • Neural networksβ˜…β˜…β˜…β˜…β˜…frequentAI and machine learning

    Xavier/He initialization chooses weight variances so activations keep a stable variance through layers.

Where is it used?

Computing topics reachable from here, through the chain of ideas that leads to them:

βš™ Robotics and control

β„’ AI and machine learning

This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.

↑ ↓ to navigate Β· ↡ Β· Esc