Generative models

Level SpecializationDifficulty ★★★★★Application⌖ Open in the map

What is it?

Models that learn a probability density of data and sample from it. Calculus is everywhere: change of variables and Jacobian determinants (flows), ELBO integrals (VAEs), Lipschitz critics (WGANs), and stochastic/ordinary differential equations integrated backwards in time (diffusion).

Formulas

dx=[f(x,t)−g(t)2 ∇xlog⁡pt(x)] dt+g(t) dwˉ\dd x = \big[f(x,t) - g(t)^2\,\nabla_x\log p_t(x)\big]\,\dd t + g(t)\,\dd\bar w
reverse-time SDE of score-based diffusion
log⁡pX(x)=log⁡pZ(f−1(x))+log⁡∣det⁡Jf−1(x)∣\log p_X(x) = \log p_Z\big(f^{-1}(x)\big) + \log\big|\det J_{f^{-1}}(x)\big|
normalizing flow likelihood

The mathematics behind it

  • Jacobian matrix★★★★★fundamental

    Normalizing flows compute exact likelihoods via log⁡pX=log⁡pZ+log⁡∣det⁡J∣\log p_X = \log p_Z + \log|\det J|, designing layers with cheap determinants.

  • Probability density function★★★★★fundamental

    Generative models learn densities; diffusion models learn the score ∇xlog⁡pt(x)\nabla_x\log p_t(x).

  • Maximum likelihood estimation★★★★★fundamental

    Autoregressive models and flows maximize likelihood exactly; VAEs and diffusion models maximize a lower bound (ELBO).

  • Inverse functions★★★★★advanced

    Normalizing flows are invertible networks: sample with ff, evaluate densities with f−1f^{-1}.

  • Integration by substitution★★★★★advanced

    In 1D, a normalizing flow is exactly the density change-of-variables formula.

  • Likelihoods of continuous generative models are multiple integrals; flows use the change-of-variables formula.

  • Continuous random variables★★★★★frequent

    Generative models are continuous random variables whose distribution imitates the data.

  • Continuous distributions★★★★★frequent

    Diffusion models add Gaussian noise; VAEs use Gaussian latent variables.

  • Divergence★★★★★advanced

    Continuous normalizing flows track log-density with the divergence of the velocity field (instantaneous change of variables).

  • Euler's method★★★★★advanced

    Diffusion and flow-matching samplers integrate a learned ODE/SDE with Euler-type steps.

  • Lipschitz continuity★★★★★advanced

    Wasserstein GANs require a 1-Lipschitz critic, enforced by weight clipping, gradient penalties or spectral normalization.

  • Partial differential equations★★★★★advanced

    Diffusion models are tied to the Fokker–Planck PDE of a noising process, reversed by learning the score ∇log⁡p\nabla\log p.

  • Equilibria and stability★★★★★advanced

    GAN training dynamics can cycle around equilibria instead of converging; stabilization tricks target this.

This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.

↑ ↓ to navigate · ↵ · Esc