Regularization

Level AdvancedDifficulty ★★★★★Application⌖ Open in the map

What is it?

Add a penalty to the loss, L+λ∥θ∥2L + \lambda\norm\theta^2 (weight decay) or λ∥θ∥1\lambda\norm\theta_1 (sparsity), to prefer simpler models and reduce overfitting. Its gradient 2λθ2\lambda\theta shrinks every weight a little each step.

Formulas

∇(L+λ∥θ∥2)=∇L+2λθ\nabla\big(L + \lambda\norm\theta^2\big) = \nabla L + 2\lambda\theta
wridge∗=(X𝖳X+λI)−1X𝖳yw^\ast_{\text{ridge}} = (X^{\mathsf T}X + \lambda I)^{-1}X^{\mathsf T}y
ridge regression also cures ill-conditioning

The mathematics behind it

  • Lagrange multipliers★★★★★advanced

    Penalized training L+λ∥w∥2L + \lambda\norm w^2 is the Lagrangian of training under a norm constraint ∥w∥≤r\norm w \le r.

Where is it used?

Computing topics reachable from here, through the chain of ideas that leads to them:

What depends on it

This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.

↑ ↓ to navigate · ↵ · Esc