Maximum likelihood estimation

Level AdvancedDifficulty ★★★★★Concept⌖ Open in the map

What is it?

Choose the parameters that make the observed data most probable: maximize ∏ip(xi∣θ)\prod_i p(x_i\mid\theta), i.e. minimize −∑ilog⁡p(xi∣θ)-\sum_i\log p(x_i\mid\theta). Squared error, cross-entropy and the training objective of language models are all negative log-likelihoods.

Why does it exist?

We need a principled way to fit a model to data. Fisher's answer (1922): a parameter value is plausible to the extent that it predicted what actually happened. Under mild conditions the MLE is consistent and asymptotically the most efficient estimator.

Intuition

Slide a Gaussian left and right over a set of points: the position where the product of the heights at the points is largest is the sample mean. Taking logs turns the product into a sum (numerically stable, easy to differentiate), and setting the derivative to zero finds the peak.

Formal definition

Given i.i.d. data x1,…,xNx_1, \dots, x_N and a model p(x∣θ)p(x\mid\theta),

θ^MLE=arg⁡max⁡θ∑i=1Nlog⁡p(xi∣θ)=arg⁡min⁡θ(−1N∑i=1Nlog⁡p(xi∣θ)).\hat\theta_{\text{MLE}} = \arg\max_\theta\sum_{i=1}^N\log p(x_i\mid\theta) = \arg\min_\theta\Big(-\frac1N\sum_{i=1}^N\log p(x_i\mid\theta)\Big).

Typically found by solving ∇θ∑ilog⁡p(xi∣θ)=0\nabla_\theta\sum_i\log p(x_i\mid\theta) = 0, in closed form or by gradient methods.

Formulas

θ^=arg⁡max⁡θ∑ilog⁡p(xi∣θ)\hat\theta = \arg\max_\theta \sum_i \log p(x_i\mid\theta)
y=fθ(x)+ε, ε∼𝒩(0,σ2)  ⟹  −log⁡p=(y−fθ(x))22σ2+consty = f_\theta(x) + \varepsilon,\ \varepsilon\sim\mathcal N(0,\sigma^2) \implies -\log p = \frac{(y - f_\theta(x))^2}{2\sigma^2} + \text{const}
Gaussian noise ⇒ squared error
−∑tlog⁡pθ(wt∣w<t)-\sum_t \log p_\theta(w_t \mid w_{<t})
the training loss of a language model

Example

Gaussian with unknown mean, known σ\sigma: ℓ(μ)=−∑(xi−μ)22σ2+const\ell(\mu) = -\sum\frac{(x_i - \mu)^2}{2\sigma^2} + \text{const}; ℓ′(μ)=∑xi−μσ2=0⇒μ^=xˉ\ell'(\mu) = \sum\frac{x_i - \mu}{\sigma^2} = 0 \Rightarrow \hat\mu = \bar x. Coin with kk heads in nn tosses: ℓ(p)=kln⁡p+(n−k)ln⁡(1−p)\ell(p) = k\ln p + (n - k)\ln(1 - p), ℓ′(p)=0⇒p^=k/n\ell'(p) = 0 \Rightarrow \hat p = k/n.

Why does it matter?

Most loss functions of machine learning are maximum likelihood in disguise, which tells you which loss matches which noise assumption. Logistic regression, softmax classifiers, language models and normalizing flows are trained by maximum likelihood; diffusion models by a bound on it.

Where it shows up in AI

  • Loss function★★★★★fundamentalAI and machine learning

    MSE and cross-entropy are negative log-likelihoods under Gaussian and categorical models.

  • Logistic regression★★★★★fundamentalAI and machine learning

    Logistic regression is the MLE of a Bernoulli model with p=σ(w⋅x+b)p = \sigma(w\cdot x + b).

  • Generative models★★★★★fundamentalAI and machine learning

    Autoregressive models and flows maximize likelihood exactly; VAEs and diffusion models maximize a lower bound (ELBO).

Where is it used?

Computing topics reachable from here, through the chain of ideas that leads to them:

Exercises

1Computation

Find the MLE of λ\lambda for exponential data x1,…,xNx_1, \dots, x_N.

Solution

ℓ(λ)=Nln⁡λ−λ∑xi\ell(\lambda) = N\ln\lambda - \lambda\sum x_i; ℓ′(λ)=N/λ−∑xi=0⇒λ^=1/xˉ\ell'(\lambda) = N/\lambda - \sum x_i = 0 \Rightarrow \hat\lambda = 1/\bar x.

2AI

Show that if the noise is Laplace, p(ε)∝e−∣ε∣/bp(\varepsilon) \propto e^{-|\varepsilon|/b}, maximum likelihood regression minimizes the absolute error.

Solution

−log⁡p(y∣x)=∣y−fθ(x)∣/b+const-\log p(y\mid x) = |y - f_\theta(x)|/b + \text{const}; summing over the data, maximizing likelihood is minimizing ∑∣yi−fθ(xi)∣\sum|y_i - f_\theta(x_i)|.

↑ ↓ to navigate · ↵ · Esc