Extrema in several variables

Level UniversityDifficulty ★★★★★Concept⌖ Open in the map

What is it?

Solve ∇f=0\nabla f = 0, then classify with the Hessian. For least squares ∥Xw−y∥2\norm{Xw - y}^2 this gives the normal equations X𝖳Xw=X𝖳yX^{\mathsf T}Xw = X^{\mathsf T}y — linear regression in closed form.

Formulas

∇w∥Xw−y∥2=2X𝖳(Xw−y)=0  ⟺  X𝖳X w=X𝖳y\nabla_w \norm{Xw - y}^2 = 2X^{\mathsf T}(Xw - y) = 0 \iff X^{\mathsf T}X\,w = X^{\mathsf T}y

Where it shows up in AI

  • Linear regression★★★★★fundamentalAI and machine learning

    Ordinary least squares is solved by setting the gradient to zero: the normal equations.

  • Loss landscape★★★★★frequentAI and machine learning

    In high dimension most critical points of a random-looking loss are saddles, not minima.

Where is it used?

Computing topics reachable from here, through the chain of ideas that leads to them:

This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.

↑ ↓ to navigate · ↵ · Esc