Probability and calculus
Continuous probability is calculus: densities are integrated to get probabilities, expectations are integrals, and maximum likelihood is setting a derivative to zero. The mathematical core of statistics and probabilistic AI.
7 topics
For discrete outcomes probability is counting; for continuous ones (a temperature, a pixel value, the weights of a model) single values have probability zero and everything is expressed with densities and integrals. That is where integrals enter machine learning: expected losses, likelihoods, Bayesian posteriors and the noise schedules of diffusion models. For the discrete and information-theoretic side, see Math of AI · Probability.
Topics
Continuous random variables
A random quantity that can take any value in an interval. for every single ; only intervals have positive probability, given by integrating a density.
Probability density function
A function with whose integral over a set is the probability of that set. A density is not a probability: it can exceed 1; it is probability per unit length.
Cumulative distribution function
. By the fundamental theorem, . Its inverse (the quantile function) turns uniform random numbers into samples of any distribution.
Expectation
: the probability-weighted average, the centre of mass of the distribution. More generally — and the expected loss over the data distribution is what learning really minimizes.
Variance
: the average squared distance from the mean. The variance of an average of independent samples is — the law that governs Monte Carlo error, batch sizes and A/B tests.
Continuous distributions
The workhorses: uniform (random number generators), exponential (waiting times, memoryless), normal (sums of many small effects, by the central limit theorem), and their multivariate versions.
Maximum likelihood estimation
Choose the parameters that make the observed data most probable: maximize , i.e. minimize . Squared error, cross-entropy and the training objective of language models are all negative log-likelihoods.
Where this area leads in computing
ℒ AI and machine learning ★★★★★
- Loss function★★★★★←Probability density function, Expectation, Maximum likelihood estimation
- Logistic regression★★★★★←Maximum likelihood estimation
- Stochastic gradient descent (SGD)★★★★★←Expectation, Variance
- Reinforcement learning★★★★★←Expectation
- Bayesian inference★★★★★←Probability density function
- Generative models★★★★★←Continuous random variables, Probability density function, Continuous distributions, Maximum likelihood estimation
- Momentum and Adam★★★★★←Variance
- Neural networks★★★★★←Variance