Bayesian inference

Level AdvancedDifficulty ★★★★★Application⌖ Open in the map

What is it?

Treat parameters as random and update a prior density into a posterior with Bayes' rule. The normalizing integral is intractable in general, so practice uses MCMC (sampling) or variational inference (optimization).

Formulas

p(θ∣D)=p(D∣θ) p(θ)∫p(D∣θ′) p(θ′) dθ′p(\theta\mid D) = \frac{p(D\mid\theta)\,p(\theta)}{\int p(D\mid\theta')\,p(\theta')\,\dd\theta'}
log⁡p(D)≥𝔼q[log⁡p(D,θ)−log⁡q(θ)]\log p(D) \ge \E_{q}\big[\log p(D,\theta) - \log q(\theta)\big]
evidence lower bound (ELBO)

The mathematics behind it

  • Probability density function★★★★★fundamental

    Priors, likelihoods and posteriors over continuous parameters are densities.

  • Improper integrals★★★★★fundamental

    The evidence p(D)=∫p(D∣θ) p(θ) dθp(D) = \int p(D\mid\theta)\,p(\theta)\,\dd\theta is an integral over the whole parameter space.

  • Numerical integration (quadrature)★★★★★frequent

    Low-dimensional posteriors and marginal likelihoods can be computed by quadrature.

Where is it used?

Computing topics reachable from here, through the chain of ideas that leads to them:

What depends on it

This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.

↑ ↓ to navigate · ↵ · Esc