What is it?
At a constrained optimum of subject to , the gradients are parallel: . The multiplier measures how much the optimum would improve if the constraint were relaxed.
Why does it exist?
Substituting the constraint to eliminate a variable works only in easy cases. Lagrange's method keeps all the variables and adds one unknown per constraint, turning a constrained problem into an unconstrained system of equations — and the new unknowns turn out to carry economic and physical meaning.
Intuition
Walk along the curve looking at the level sets of . While the curve crosses level sets, you can still go lower. At the optimum the curve is tangent to a level set of , so their normals — and — point along the same line.
Formal definition
If is a local extremum of on , with continuously differentiable and , then there is with . Equivalently, is a critical point of the Lagrangian
Formulas
- the multiplier is the sensitivity of the optimum ("shadow price")
Example
Maximize entropy subject to . , so all are equal: the uniform distribution. Add a constraint on the expected energy and the same computation produces the softmax / Boltzmann distribution .
Why does it matter?
Support vector machines are derived through their Lagrangian dual (where the kernel trick appears). Constrained training, fairness constraints, the maximum-entropy justification of softmax, and duality in operations research all use the same idea. In physics, Lagrange's other great idea — the Lagrangian of mechanics — runs robot dynamics.
Where it shows up in computing
Dual variables are shadow prices: how much an extra unit of a resource is worth.
Where it shows up in AI
The SVM dual is obtained with Lagrange multipliers; only points with (support vectors) matter.
Penalized training is the Lagrangian of training under a norm constraint .
Where is it used?
Computing topics reachable from here, through the chain of ideas that leads to them:
What depends on it
Exercises
Maximize subject to with a Lagrange multiplier.
Solution
gives , and gives , , .
Show that the distribution maximizing entropy with a fixed mean energy has the form .
Solution
, so : a softmax of .