Loss function

Level UniversityDifficulty ★★★★★Application⌖ Open in the map

What is it?

A single number L(θ)L(\theta) measuring how wrong a model with parameters θ\theta is on the data. Learning is minimizing it. Mean squared error for regression and cross-entropy for classification are the two everyone uses.

Why does it exist?

"Make the model good" is not something an algorithm can optimize; "make this differentiable number small" is. The loss converts the goal into a function of the parameters, so calculus — derivatives, gradients, convexity — can take over.

Intuition

A landscape over parameter space: height = error. Each point is a possible model; training walks downhill. The shape of the loss decides what "wrong" means: squared error punishes large mistakes heavily, absolute error is robust to outliers, cross-entropy punishes confident wrong probabilities without limit.

Formal definition

Empirical risk over a dataset {(xi,yi)}i=1N\{(x_i, y_i)\}_{i=1}^N:

L(θ)=1N∑i=1Nℓ(fθ(xi),yi),L(\theta) = \frac1N\sum_{i=1}^N \ell\big(f_\theta(x_i), y_i\big),

an estimate of the expected risk 𝔼(x,y)∼P[ℓ(fθ(x),y)]\E_{(x,y)\sim P}[\ell(f_\theta(x), y)].

Formulas

MSE=1N∑i(y^i−yi)2\text{MSE} = \frac1N\sum_i\big(\hat y_i - y_i\big)^2
CE=−1N∑i∑kyiklog⁡p^ik,p^=softmax⁡(z)\text{CE} = -\frac1N\sum_i\sum_k y_{ik}\log\hat p_{ik}, \qquad \hat p = \operatorname{softmax}(z)
∂ CE∂zk=p^k−yk\frac{\partial\,\text{CE}}{\partial z_k} = \hat p_k - y_k
softmax + cross-entropy: the gradient is simply "prediction minus target"

Why does it matter?

Choosing the loss is choosing what the model will be good at. It must be differentiable (almost everywhere) for gradient training, and numerically stable — which is why frameworks fuse softmax and cross-entropy into one operation.

The mathematics behind it

  • Logarithmic functions★★★★★fundamental

    Cross-entropy is a negative log-likelihood; the log turns a product over samples into a sum.

  • Maxima and minima★★★★★fundamental

    Training is the search for parameters that minimize the loss — in practice, a good local minimum.

  • Convexity and concavity★★★★★fundamental

    Squared error and cross-entropy are convex in the predictions; for linear models, in the parameters too.

  • Functions of several variables★★★★★fundamental

    A loss is a scalar function of all the model's parameters.

  • Expectation★★★★★fundamental

    The objective of learning is the expected loss (risk) over the data distribution.

  • Maximum likelihood estimation★★★★★fundamental

    MSE and cross-entropy are negative log-likelihoods under Gaussian and categorical models.

  • Numerical stability★★★★★frequent

    Libraries fuse softmax and cross-entropy (log-sum-exp) because the naive composition overflows.

  • Probability density function★★★★★fundamental

    Regression losses are negative log-densities: squared error ↔ Gaussian noise, absolute error ↔ Laplace noise.

  • Absolute value★★★★★frequent

    The L1 / mean absolute error loss 1n∑∣yi−y^i∣\frac1n\sum|y_i - \hat y_i| is robust to outliers but not differentiable at 0.

  • L1 and hinge losses are non-differentiable at a point and are optimized with subgradients.

  • Definite integral★★★★★frequent

    The true risk 𝔼[ℓ]=∫ℓ dP\E[\ell] = \int \ell\,\dd P is an integral; the training loss is its sample average.

Where is it used?

Computing topics reachable from here, through the chain of ideas that leads to them:

What depends on it

Exercises

1AI

A classifier outputs p^=0.01\hat p = 0.01 for the true class. What is its cross-entropy? And with p^=0.9\hat p = 0.9?

Solution

−ln⁡0.01≈4.6-\ln 0.01 \approx 4.6 vs −ln⁡0.9≈0.105-\ln 0.9 \approx 0.105. Confidently wrong predictions are punished ~44 times more.

↑ ↓ to navigate · ↵ · Esc