∇ Mathematics for AI

The shortest honest path from derivatives to deep learning: gradients, optimization, likelihood and backpropagation.

0/14
  1. 01

    Functions

    A rule ff that assigns to each input xx of a set AA exactly one output f(x)f(x) in a set BB. Calculus studies how outputs change when inputs change; computing is, quite literally, evaluating functions.

    FundamentalFoundations
  2. 02

    Derivative

    f′(a)f'(a) is the instantaneous rate of change of ff at aa: the slope of the tangent line to the graph, defined as the limit of slopes of secant lines.

    FundamentalDerivatives
  3. 03

    Chain rule

    The derivative of a composition is the product of the derivatives: (g∘f)′(x)=g′(f(x)) f′(x)(g\circ f)'(x) = g'(f(x))\,f'(x). Rates of change multiply along a chain — and backpropagation is this rule applied, very efficiently, to a neural network.

    FundamentalDerivatives
  4. 04

    Partial derivatives

    ∂f∂xi\frac{\partial f}{\partial x_i}: the derivative with respect to one variable, holding the others fixed. Each answers "how sensitive is the output to this input?" — for a neural network, to this weight.

    UniversityMultivariable calculus
  5. 05

    Gradient

    ∇f=(∂f∂x1,…,∂f∂xn)\nabla f = \left(\frac{\partial f}{\partial x_1}, \dots, \frac{\partial f}{\partial x_n}\right): the vector of all partial derivatives. It points in the direction of steepest ascent, its length is that steepest slope, and it is perpendicular to the level sets. Walk against it and you go downhill fastest.

    UniversityMultivariable calculus
  6. 06

    Multivariable chain rule

    When a variable influences the output through several paths, add the contributions of every path, each the product of the local derivatives along it: ∂z∂x=∑i∂z∂ui∂ui∂x\frac{\partial z}{\partial x} = \sum_i \frac{\partial z}{\partial u_i}\frac{\partial u_i}{\partial x}. In matrix form, Jacobians multiply. This is exactly what backpropagation computes on a network's graph.

    UniversityMultivariable calculus
  7. 07

    Convexity and concavity

    ff is convex if the chord between any two points of its graph lies above the graph — equivalently, for smooth ff, if f′′≥0f'' \ge 0. For convex functions every local minimum is global, which is why convex problems are the ones optimization can reliably solve.

    UniversityExtrema and optimization
  8. 08

    Gradient descent

    Repeat θ←θ−η ∇L(θ)\theta \leftarrow \theta - \eta\,\nabla L(\theta): take a small step against the gradient. Cauchy proposed it in 1847; today it (and its stochastic, adaptive variants) trains essentially every neural network.

    UniversityAI and machine learning
  9. 09

    Probability density function

    A function p(x)≥0p(x) \ge 0 with ∫p=1\int p = 1 whose integral over a set is the probability of that set. A density is not a probability: it can exceed 1; it is probability per unit length.

    UniversityProbability and calculus
  10. 10

    Maximum likelihood estimation

    Choose the parameters that make the observed data most probable: maximize ∏ip(xi∣θ)\prod_i p(x_i\mid\theta), i.e. minimize −∑ilog⁡p(xi∣θ)-\sum_i\log p(x_i\mid\theta). Squared error, cross-entropy and the training objective of language models are all negative log-likelihoods.

    AdvancedProbability and calculus
  11. 11

    Loss function

    A single number L(θ)L(\theta) measuring how wrong a model with parameters θ\theta is on the data. Learning is minimizing it. Mean squared error for regression and cross-entropy for classification are the two everyone uses.

    UniversityAI and machine learning
  12. 12

    Neural networks

    Compositions of affine maps and non-linearities, fθ=fL∘⋯∘f1f_\theta = f_L\circ\dots\circ f_1 with fℓ(a)=σ(Wℓa+bℓ)f_\ell(a) = \sigma(W_\ell a + b_\ell). A differentiable function with millions of adjustable parameters, trained by gradient descent.

    AdvancedAI and machine learning
  13. 13

    Backpropagation

    The algorithm that computes ∂L/∂W\partial L/\partial W and ∂L/∂b\partial L/\partial b for every layer of a network: one forward pass storing intermediate values, then one backward pass applying the chain rule from the loss down to the inputs. Cost: about twice the forward pass, whatever the number of parameters.

    AdvancedAI and machine learning
  14. 14

    Deep learning

    Training very deep networks (CNNs, transformers) with backpropagation and adaptive SGD on huge datasets. The calculus is the same as for one neuron; what changed is scale, architecture (residual connections, normalization, attention) and hardware.

    AdvancedAI and machine learning
↑ ↓ to navigate · ↵ · Esc