Probability density function

Level UniversityDifficulty ★★★★★Concept⌖ Open in the map

What is it?

A function p(x)≥0p(x) \ge 0 with ∫p=1\int p = 1 whose integral over a set is the probability of that set. A density is not a probability: it can exceed 1; it is probability per unit length.

Why does it exist?

With infinitely many possible values, each individual value must have probability zero, yet some regions are more likely than others. The density describes that relative likelihood, and integration recovers actual probabilities.

Intuition

Think of a histogram of many samples with bins getting thinner and thinner, normalized so the total area is 1. In the limit the bars become a curve: the density. Probability is area under it; p(x) dxp(x)\,\dd x is the probability of landing in a tiny interval near xx.

Formal definition

XX has density pp if P(X∈A)=∫Ap(x) dxP(X \in A) = \int_A p(x)\,\dd x for every (measurable) set AA, with p≥0p \ge 0 and ∫−∞∞p=1\int_{-\infty}^{\infty}p = 1. Then p=F′p = F' where FF is the cumulative distribution function.

Formulas

p(x)≥0,∫−∞∞p(x) dx=1p(x) \ge 0, \qquad \int_{-\infty}^{\infty} p(x)\,\dd x = 1
𝒩(x;μ,σ2)=1σ2πexp⁡ ⁣(−(x−μ)22σ2)\mathcal N(x;\mu,\sigma^2) = \frac{1}{\sigma\sqrt{2\pi}}\exp\!\Big(-\frac{(x - \mu)^2}{2\sigma^2}\Big)
normal (Gaussian) density
p(θ∣D)=p(D∣θ) p(θ)∫p(D∣θ′) p(θ′) dθ′p(\theta\mid D) = \frac{p(D\mid\theta)\,p(\theta)}{\int p(D\mid\theta')\,p(\theta')\,\dd\theta'}
Bayes' rule for densities

Example

Uniform on [0,0.5][0, 0.5]: p(x)=2p(x) = 2 there. The density is 2 (> 1) but every probability is fine: P(X≤0.1)=∫00.12 dx=0.2P(X \le 0.1) = \int_0^{0.1}2\,\dd x = 0.2.

Why does it matter?

Generative AI is density modelling: VAEs, normalizing flows and diffusion models all learn a density pθ(x)p_\theta(x) of images or sounds (explicitly or through its score ∇xlog⁡p\nabla_x\log p). Bayesian inference updates densities with data, and the normalizing integral in the denominator is usually the hard part.

Where it shows up in computing

  • Kalman filter★★★★★frequentRobotics and control

    The Kalman filter propagates Gaussian densities of the state through predictions and measurements.

Where it shows up in AI

  • Generative models★★★★★fundamentalAI and machine learning

    Generative models learn densities; diffusion models learn the score ∇xlog⁡pt(x)\nabla_x\log p_t(x).

  • Bayesian inference★★★★★fundamentalAI and machine learning

    Priors, likelihoods and posteriors over continuous parameters are densities.

  • Loss function★★★★★fundamentalAI and machine learning

    Regression losses are negative log-densities: squared error ↔ Gaussian noise, absolute error ↔ Laplace noise.

Where is it used?

Computing topics reachable from here, through the chain of ideas that leads to them:

What depends on it

Exercises

1Computation

For which cc is p(x)=c x(1−x)p(x) = c\,x(1 - x) on [0,1][0,1] a density? Compute P(X>0.5)P(X > 0.5).

Solution

∫01x(1−x) dx=16\int_0^1 x(1-x)\,\dd x = \frac16, so c=6c = 6. By symmetry P(X>0.5)=0.5P(X > 0.5) = 0.5.

↑ ↓ to navigate · ↵ · Esc