What is it?
Models that learn a probability density of data and sample from it. Calculus is everywhere: change of variables and Jacobian determinants (flows), ELBO integrals (VAEs), Lipschitz critics (WGANs), and stochastic/ordinary differential equations integrated backwards in time (diffusion).
Formulas
- reverse-time SDE of score-based diffusion
- normalizing flow likelihood
The mathematics behind it
Normalizing flows compute exact likelihoods via , designing layers with cheap determinants.
Generative models learn densities; diffusion models learn the score .
Autoregressive models and flows maximize likelihood exactly; VAEs and diffusion models maximize a lower bound (ELBO).
Normalizing flows are invertible networks: sample with , evaluate densities with .
In 1D, a normalizing flow is exactly the density change-of-variables formula.
Likelihoods of continuous generative models are multiple integrals; flows use the change-of-variables formula.
Generative models are continuous random variables whose distribution imitates the data.
Diffusion models add Gaussian noise; VAEs use Gaussian latent variables.
Continuous normalizing flows track log-density with the divergence of the velocity field (instantaneous change of variables).
Diffusion and flow-matching samplers integrate a learned ODE/SDE with Euler-type steps.
Wasserstein GANs require a 1-Lipschitz critic, enforced by weight clipping, gradient penalties or spectral normalization.
Diffusion models are tied to the Fokker–Planck PDE of a noising process, reversed by learning the score .
GAN training dynamics can cycle around equilibria instead of converging; stabilization tricks target this.
This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.