Critical points and derivative tests

Level FundamentalDifficulty ★★★★★Concept⌖ Open in the map

What is it?

Where f′(x)=0f'(x) = 0 (or does not exist). Fermat: an interior extremum of a differentiable function is a critical point. The sign of f′′f'' then tells minimum (f′′>0f'' > 0), maximum (f′′<0f'' < 0) or "look closer" (f′′=0f'' = 0).

Why does it exist?

Checking every point to find the lowest is impossible on a continuum. Fermat's observation reduces the search to solving an equation, f′(x)=0f'(x) = 0, usually with finitely many solutions. Optimization algorithms are, at heart, methods to solve that equation when it cannot be solved by hand.

Intuition

At the bottom of a valley or the top of a hill the tangent is horizontal. But a horizontal tangent can also be an inflection point (x3x^3 at 0) — and in many dimensions, a saddle: a minimum in one direction and a maximum in another. In high-dimensional loss landscapes most critical points are saddles.

Formal definition

Fermat. If ff has a local extremum at an interior point cc and f′(c)f'(c) exists, then f′(c)=0f'(c) = 0.

Second derivative test. If f′(c)=0f'(c) = 0 and f′′(c)>0f''(c) > 0, cc is a strict local minimum; if f′′(c)<0f''(c) < 0, a strict local maximum.

Formulas

f′(c)=0,f′′(c)>0  ⟹  local minimumf'(c) = 0, \quad f''(c) > 0 \implies \text{local minimum}
min⁡w∑i(yi−wxi)2  ⟹  w∗=∑ixiyi∑ixi2\min_w \sum_i (y_i - w x_i)^2 \implies w^\ast = \frac{\sum_i x_i y_i}{\sum_i x_i^2}
least squares through the origin: set the derivative to zero

How is it computed?

  1. Compute f′f' and solve f′(x)=0f'(x) = 0; add points where f′f' does not exist.
  2. Classify each with the sign of f′′f'' (or the sign change of f′f').
  3. For global extrema on [a,b][a,b], also compare the endpoint values. When step 1 has no closed-form solution, iterate: gradient descent or Newton's method on f′f'.

Example

Fit y≈wxy \approx w x to the points (1,2),(2,3),(3,7)(1, 2), (2, 3), (3, 7). Loss L(w)=(2−w)2+(3−2w)2+(7−3w)2L(w) = (2 - w)^2 + (3 - 2w)^2 + (7 - 3w)^2. L′(w)=−2(2−w)−4(3−2w)−6(7−3w)=28w−58=0⇒w=58/28≈2.07L'(w) = -2(2 - w) - 4(3 - 2w) - 6(7 - 3w) = 28w - 58 = 0 \Rightarrow w = 58/28 \approx 2.07, and L′′=28>0L'' = 28 > 0: a minimum. This is linear regression in one line.

Why does it matter?

"Set the derivative to zero" is the most reused recipe of applied mathematics: least squares, maximum likelihood, optimal control and economics all start there. Gradient-based training is that recipe executed numerically when the equation is too big to solve.

Where it shows up in AI

  • Gradient descent★★★★★fundamentalAI and machine learning

    Gradient descent stops where the derivative vanishes — at a critical point.

  • Loss landscape★★★★★fundamentalAI and machine learning

    Minima, maxima and (overwhelmingly) saddle points shape the landscape optimizers must cross.

Where is it used?

Computing topics reachable from here, through the chain of ideas that leads to them:

What depends on it

Exercises

1Computation

Find and classify the critical points of f(x)=x4−4x3f(x) = x^4 - 4x^3.

Solution

f′(x)=4x2(x−3)f'(x) = 4x^2(x - 3): critical points 0 and 3. f′(x)=12x2−24xf'(x) = 12x^2 - 24x; f′(3)=36>0f'(3) = 36 > 0 → minimum. At 0, f′=0f' = 0 and f′f' does not change sign (negative on both sides): not an extremum.

2Applied

A server costs c(n)=100/n+4nc(n) = 100/n + 4n (latency penalty plus hardware) with nn instances. Which nn minimizes cost?

Solution

c′(n)=−100/n2+4=0⇒n=5c'(n) = -100/n^2 + 4 = 0 \Rightarrow n = 5; c′(n)=200/n3>0c'(n) = 200/n^3 > 0, so it is a minimum, c(5)=40c(5) = 40.

↑ ↓ to navigate · ↵ · Esc