What is it?
Compositions of affine maps and non-linearities, with . A differentiable function with millions of adjustable parameters, trained by gradient descent.
Why does it exist?
Linear models cannot represent XOR, let alone images or language. Stacking simple differentiable layers gives a family of functions flexible enough to approximate almost anything (universal approximation) while keeping the property that matters most: we can compute the gradient of the loss with respect to every parameter.
Intuition
Each layer bends and folds the space a little: a linear map rotates and stretches, the activation clips or squashes. After enough folds, classes that were tangled become separable by a plane. A single neuron is a soft step across a hyperplane; sums of many such steps can build any bump.
Formal definition
Universal approximation (Cybenko 1989, Hornik 1991): one hidden layer with enough neurons and a non-polynomial activation approximates any continuous function on a compact set to any accuracy.
Formulas
- one layer
- a 3-layer perceptron
Why does it matter?
Vision, speech, translation, protein folding, game playing and language models are neural networks. Mathematically they are "just" compositions of differentiable functions — which is exactly why calculus can train them.
The mathematics behind it
A network is a parametrized function ; training chooses .
A deep network is the composition of its layers, .
Inputs, activations and parameters are vectors (and tensors) in .
Each neuron computes before its activation.
A dense layer is ; GPUs exist to multiply these matrices fast.
Hopfield networks store memories as point attractors of an energy-descending dynamics (Nobel Prize in Physics 2024).
Xavier/He initialization chooses weight variances so activations keep a stable variance through layers.
Recurrent networks are discrete dynamical systems ; exploding and vanishing gradients are questions of stability.
Proofs of the universal approximation theorem use uniform continuity on compact sets.
Where is it used?
Computing topics reachable from here, through the chain of ideas that leads to them: