What is it?
: the probability-weighted average, the centre of mass of the distribution. More generally β and the expected loss over the data distribution is what learning really minimizes.
Why does it exist?
To summarize a random quantity by one number that behaves well: it is linear, it is what averages of many samples converge to (law of large numbers), and decisions that maximize expected value are optimal in the long run.
Intuition
Balance the graph of the density on a knife edge: it balances at . And the average of independent samples estimates it with error about β the principle behind Monte Carlo, mini-batch gradients and A/B tests.
Formal definition
when the integral converges absolutely. Linearity: (always, even for dependent variables). Law of large numbers: for i.i.d. samples.
Formulas
- expected risk and its mini-batch estimate
- why stochastic gradients are unbiased
Why does it matter?
Training minimizes an expectation it cannot compute, by following unbiased noisy estimates of its gradient (SGD). Reinforcement learning maximizes expected return. Monte Carlo rendering computes pixel colours as expectations over random light paths.
Where it shows up in computing
Monte Carlo estimates expectations by sample averages.
Little's law relates expected queue length and expected waiting time.
Where it shows up in AI
The objective of learning is the expected loss (risk) over the data distribution.
A mini-batch gradient is an unbiased estimate of the expected gradient.
Agents maximize expected discounted return; policy gradients differentiate an expectation.
Where is it used?
Computing topics reachable from here, through the chain of ideas that leads to them:
β AI and machine learning
- Loss functionβ β β β β
- Stochastic gradient descent (SGD)β β β β β
- Reinforcement learningβ β β β β
- Loss functionβGradient descentβ β β β β
- Loss functionβLogistic regressionβ β β β β
- Stochastic gradient descent (SGD)βMomentum and Adamβ β β β β
- +11
3D Computer graphics
- Monte Carlo methodsβThe rendering equationβ β β β β
β Robotics and control
- VarianceβKalman filterβ β β β β
β Optimization and systems
- Loss functionβGradient descentβRecommender systemsβ β β β β
- Queueing theory and performanceβ β β β β
What depends on it
Exercises
Compute for the exponential density , .
Solution
By parts: .
Why is the gradient of a random mini-batch loss an unbiased estimate of the full gradient? What assumption is needed?
Solution
If the batch is drawn uniformly at random, by linearity of expectation (and exchanging gradient and expectation, which needs mild smoothness). Non-random batches (e.g. sorted data) break it.