What is it?
Choose the parameters that make the observed data most probable: maximize , i.e. minimize . Squared error, cross-entropy and the training objective of language models are all negative log-likelihoods.
Why does it exist?
We need a principled way to fit a model to data. Fisher's answer (1922): a parameter value is plausible to the extent that it predicted what actually happened. Under mild conditions the MLE is consistent and asymptotically the most efficient estimator.
Intuition
Slide a Gaussian left and right over a set of points: the position where the product of the heights at the points is largest is the sample mean. Taking logs turns the product into a sum (numerically stable, easy to differentiate), and setting the derivative to zero finds the peak.
Formal definition
Given i.i.d. data and a model ,
Typically found by solving , in closed form or by gradient methods.
Formulas
- Gaussian noise ⇒ squared error
- the training loss of a language model
Example
Gaussian with unknown mean, known : ; . Coin with heads in tosses: , .
Why does it matter?
Most loss functions of machine learning are maximum likelihood in disguise, which tells you which loss matches which noise assumption. Logistic regression, softmax classifiers, language models and normalizing flows are trained by maximum likelihood; diffusion models by a bound on it.
Where it shows up in AI
MSE and cross-entropy are negative log-likelihoods under Gaussian and categorical models.
Logistic regression is the MLE of a Bernoulli model with .
Autoregressive models and flows maximize likelihood exactly; VAEs and diffusion models maximize a lower bound (ELBO).
Where is it used?
Computing topics reachable from here, through the chain of ideas that leads to them:
Exercises
Find the MLE of for exponential data .
Solution
; .
Show that if the noise is Laplace, , maximum likelihood regression minimizes the absolute error.
Solution
; summing over the data, maximizing likelihood is minimizing .