Neural networks

Level AdvancedDifficulty ★★★★★Application⌖ Open in the map

What is it?

Compositions of affine maps and non-linearities, fθ=fL∘⋯∘f1f_\theta = f_L\circ\dots\circ f_1 with fℓ(a)=σ(Wℓa+bℓ)f_\ell(a) = \sigma(W_\ell a + b_\ell). A differentiable function with millions of adjustable parameters, trained by gradient descent.

Why does it exist?

Linear models cannot represent XOR, let alone images or language. Stacking simple differentiable layers gives a family of functions flexible enough to approximate almost anything (universal approximation) while keeping the property that matters most: we can compute the gradient of the loss with respect to every parameter.

Intuition

Each layer bends and folds the space a little: a linear map rotates and stretches, the activation clips or squashes. After enough folds, classes that were tangled become separable by a plane. A single neuron σ(w⋅x+b)\sigma(w\cdot x + b) is a soft step across a hyperplane; sums of many such steps can build any bump.

Formal definition

a(0)=x,z(ℓ)=W(ℓ)a(ℓ−1)+b(ℓ),a(ℓ)=σ(z(ℓ)),y^=a(L).a^{(0)} = x, \qquad z^{(\ell)} = W^{(\ell)}a^{(\ell-1)} + b^{(\ell)}, \qquad a^{(\ell)} = \sigma\big(z^{(\ell)}\big), \qquad \hat y = a^{(L)}.

Universal approximation (Cybenko 1989, Hornik 1991): one hidden layer with enough neurons and a non-polynomial activation approximates any continuous function on a compact set to any accuracy.

Formulas

y^=σ(Wx+b)\hat y = \sigma\big(W x + b\big)
one layer
fθ(x)=W3 σ(W2 σ(W1x+b1)+b2)+b3f_\theta(x) = W_3\,\sigma\big(W_2\,\sigma(W_1 x + b_1) + b_2\big) + b_3
a 3-layer perceptron

Why does it matter?

Vision, speech, translation, protein folding, game playing and language models are neural networks. Mathematically they are "just" compositions of differentiable functions — which is exactly why calculus can train them.

The mathematics behind it

  • Functions★★★★★fundamental

    A network is a parametrized function fθ:ℝn→ℝmf_\theta : \R^n \to \R^m; training chooses θ\theta.

  • Composition★★★★★fundamental

    A deep network is the composition of its layers, fL∘⋯∘f1f_L \circ \dots \circ f_1.

  • Vectors★★★★★fundamental

    Inputs, activations and parameters are vectors (and tensors) in ℝn\R^n.

  • Dot product★★★★★fundamental

    Each neuron computes w⋅x+bw \cdot x + b before its activation.

  • Matrices and linear maps★★★★★fundamental

    A dense layer is σ(Wx+b)\sigma(Wx + b); GPUs exist to multiply these matrices fast.

  • Attractors★★★★★historical

    Hopfield networks store memories as point attractors of an energy-descending dynamics (Nobel Prize in Physics 2024).

  • Variance★★★★★frequent

    Xavier/He initialization chooses weight variances so activations keep a stable variance through layers.

  • Dynamical systems★★★★★advanced

    Recurrent networks are discrete dynamical systems ht+1=σ(Wht+Uxt)h_{t+1} = \sigma(Wh_t + Ux_t); exploding and vanishing gradients are questions of stability.

  • Uniform continuity★★★★★advanced

    Proofs of the universal approximation theorem use uniform continuity on compact sets.

Where is it used?

Computing topics reachable from here, through the chain of ideas that leads to them:

What depends on it

Further reading

↑ ↓ to navigate · ↵ · Esc