What is it?
A single number measuring how wrong a model with parameters is on the data. Learning is minimizing it. Mean squared error for regression and cross-entropy for classification are the two everyone uses.
Why does it exist?
"Make the model good" is not something an algorithm can optimize; "make this differentiable number small" is. The loss converts the goal into a function of the parameters, so calculus — derivatives, gradients, convexity — can take over.
Intuition
A landscape over parameter space: height = error. Each point is a possible model; training walks downhill. The shape of the loss decides what "wrong" means: squared error punishes large mistakes heavily, absolute error is robust to outliers, cross-entropy punishes confident wrong probabilities without limit.
Formal definition
Empirical risk over a dataset :
an estimate of the expected risk .
Formulas
- softmax + cross-entropy: the gradient is simply "prediction minus target"
Why does it matter?
Choosing the loss is choosing what the model will be good at. It must be differentiable (almost everywhere) for gradient training, and numerically stable — which is why frameworks fuse softmax and cross-entropy into one operation.
The mathematics behind it
Cross-entropy is a negative log-likelihood; the log turns a product over samples into a sum.
Training is the search for parameters that minimize the loss — in practice, a good local minimum.
Squared error and cross-entropy are convex in the predictions; for linear models, in the parameters too.
A loss is a scalar function of all the model's parameters.
The objective of learning is the expected loss (risk) over the data distribution.
MSE and cross-entropy are negative log-likelihoods under Gaussian and categorical models.
Libraries fuse softmax and cross-entropy (log-sum-exp) because the naive composition overflows.
Regression losses are negative log-densities: squared error ↔ Gaussian noise, absolute error ↔ Laplace noise.
The L1 / mean absolute error loss is robust to outliers but not differentiable at 0.
L1 and hinge losses are non-differentiable at a point and are optimized with subgradients.
The true risk is an integral; the training loss is its sample average.
Where is it used?
Computing topics reachable from here, through the chain of ideas that leads to them:
What depends on it
Exercises
A classifier outputs for the true class. What is its cross-entropy? And with ?
Solution
vs . Confidently wrong predictions are punished ~44 times more.