What is it?
Where (or does not exist). Fermat: an interior extremum of a differentiable function is a critical point. The sign of then tells minimum (), maximum () or "look closer" ().
Why does it exist?
Checking every point to find the lowest is impossible on a continuum. Fermat's observation reduces the search to solving an equation, , usually with finitely many solutions. Optimization algorithms are, at heart, methods to solve that equation when it cannot be solved by hand.
Intuition
At the bottom of a valley or the top of a hill the tangent is horizontal. But a horizontal tangent can also be an inflection point ( at 0) — and in many dimensions, a saddle: a minimum in one direction and a maximum in another. In high-dimensional loss landscapes most critical points are saddles.
Formal definition
Fermat. If has a local extremum at an interior point and exists, then .
Second derivative test. If and , is a strict local minimum; if , a strict local maximum.
Formulas
- least squares through the origin: set the derivative to zero
How is it computed?
- Compute and solve ; add points where does not exist.
- Classify each with the sign of (or the sign change of ).
- For global extrema on , also compare the endpoint values. When step 1 has no closed-form solution, iterate: gradient descent or Newton's method on .
Example
Fit to the points . Loss . , and : a minimum. This is linear regression in one line.
Why does it matter?
"Set the derivative to zero" is the most reused recipe of applied mathematics: least squares, maximum likelihood, optimal control and economics all start there. Gradient-based training is that recipe executed numerically when the equation is too big to solve.
Where it shows up in AI
Gradient descent stops where the derivative vanishes — at a critical point.
Minima, maxima and (overwhelmingly) saddle points shape the landscape optimizers must cross.
Where is it used?
Computing topics reachable from here, through the chain of ideas that leads to them:
ℒ AI and machine learning
- Extrema in several variables→Linear regression★★★★★
- Convexity and concavity→Loss function★★★★★
- Maximum likelihood estimation→Logistic regression★★★★★
- Convexity and concavity→Support vector machines★★★★★
- Maximum likelihood estimation→Generative models★★★★★
- Convexity and concavity→Loss function→Gradient descent★★★★★
- +12
What depends on it
Exercises
Find and classify the critical points of .
Solution
: critical points 0 and 3. ; → minimum. At 0, and does not change sign (negative on both sides): not an extremum.
A server costs (latency penalty plus hardware) with instances. Which minimizes cost?
Solution
; , so it is a minimum, .