Probability and calculus

Continuous probability is calculus: densities are integrated to get probabilities, expectations are integrals, and maximum likelihood is setting a derivative to zero. The mathematical core of statistics and probabilistic AI.

7 topics

For discrete outcomes probability is counting; for continuous ones (a temperature, a pixel value, the weights of a model) single values have probability zero and everything is expressed with densities and integrals. That is where integrals enter machine learning: expected losses, likelihoods, Bayesian posteriors and the noise schedules of diffusion models. For the discrete and information-theoretic side, see Math of AI · Probability.

Topics

Continuous random variables

A random quantity that can take any value in an interval. P(X=x)=0P(X = x) = 0 for every single xx; only intervals have positive probability, given by integrating a density.

University

Probability density function

A function p(x)≥0p(x) \ge 0 with ∫p=1\int p = 1 whose integral over a set is the probability of that set. A density is not a probability: it can exceed 1; it is probability per unit length.

University

Cumulative distribution function

F(x)=P(X≤x)=∫−∞xp(t) dtF(x) = P(X \le x) = \int_{-\infty}^x p(t)\,\dd t. By the fundamental theorem, F′=pF' = p. Its inverse (the quantile function) turns uniform random numbers into samples of any distribution.

University

Expectation

𝔼[X]=∫x p(x) dx\E[X] = \int x\,p(x)\,\dd x: the probability-weighted average, the centre of mass of the distribution. More generally 𝔼[g(X)]=∫g(x) p(x) dx\E[g(X)] = \int g(x)\,p(x)\,\dd x — and the expected loss over the data distribution is what learning really minimizes.

University

Variance

Var⁡(X)=𝔼[(X−𝔼X)2]\Var(X) = \E[(X - \E X)^2]: the average squared distance from the mean. The variance of an average of NN independent samples is σ2/N\sigma^2/N — the 1/N1/\sqrt N law that governs Monte Carlo error, batch sizes and A/B tests.

University

Continuous distributions

The workhorses: uniform (random number generators), exponential (waiting times, memoryless), normal (sums of many small effects, by the central limit theorem), and their multivariate versions.

University

Maximum likelihood estimation

Choose the parameters that make the observed data most probable: maximize ∏ip(xi∣θ)\prod_i p(x_i\mid\theta), i.e. minimize −∑ilog⁡p(xi∣θ)-\sum_i\log p(x_i\mid\theta). Squared error, cross-entropy and the training objective of language models are all negative log-likelihoods.

Advanced

Where this area leads in computing

↑ ↓ to navigate · ↵ · Esc