What is it?
Add a penalty to the loss, (weight decay) or (sparsity), to prefer simpler models and reduce overfitting. Its gradient shrinks every weight a little each step.
Formulas
- ridge regression also cures ill-conditioning
The mathematics behind it
Penalized training is the Lagrangian of training under a norm constraint .
Where is it used?
Computing topics reachable from here, through the chain of ideas that leads to them:
What depends on it
This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.