Lipschitz continuity

Level AdvancedDifficulty ★★★★★Concept⌖ Open in the map

What is it?

∣f(x)−f(y)∣≤L ∣x−y∣|f(x) - f(y)| \le L\,|x - y|: the function never changes faster than rate LL. When the gradient is LL-Lipschitz, gradient descent with step η<2/L\eta < 2/L is guaranteed to decrease the loss.

Formulas

∣f(x)−f(y)∣≤L ∣x−y∣|f(x) - f(y)| \le L\,|x - y|
∥∇f(x)−∇f(y)∥≤L∥x−y∥  ⟹  f(y)≤f(x)+∇f(x)⋅(y−x)+L2∥y−x∥2\norm{\nabla f(x) - \nabla f(y)} \le L\norm{x - y} \implies f(y) \le f(x) + \nabla f(x)\cdot(y - x) + \tfrac L2\norm{y - x}^2
descent lemma

Where it shows up in AI

  • Gradient descent★★★★★fundamentalAI and machine learning

    The safe learning rate is set by the Lipschitz constant of the gradient: η<2/L\eta < 2/L.

  • Generative models★★★★★advancedAI and machine learning

    Wasserstein GANs require a 1-Lipschitz critic, enforced by weight clipping, gradient penalties or spectral normalization.

Where is it used?

Computing topics reachable from here, through the chain of ideas that leads to them:

What depends on it

This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.

↑ ↓ to navigate · ↵ · Esc