Multivariable calculus

Functions of many variables: partial derivatives, the gradient, Jacobians and Hessians. The language in which a neural network with a billion parameters is trained.

13 topics

Real problems have many inputs: the pixels of an image, the joints of a robot, the weights of a network. The derivative generalizes to a vector (the gradient) or a matrix (the Jacobian), and second derivatives to a matrix (the Hessian). The single most important chain in modern computing starts here:

∇f  →  gradient descent  →  loss function  →  backpropagation  →  neural networks  →  deep learning

Topics

Functions of several variables

f:ℝn→ℝf : \R^n \to \R (or ℝm\R^m): many inputs, one (or many) outputs. With two inputs the graph is a surface over the plane; with a million inputs — the weights of a model — we reason with its level sets and its gradient.

University

Surfaces and level sets

Three ways to describe a surface: a graph z=f(x,y)z = f(x,y), a level set F(x,y,z)=cF(x,y,z) = c (implicit), or a parametrization r(u,v)r(u,v). Contour maps of loss functions and the isosurfaces of medical imaging are level sets.

University

Limits and continuity in several variables

Same ε\varepsilon–δ\delta definition with ∥x−a∥\norm{x - a} instead of ∣x−a∣|x - a|. New subtlety: the point can be approached along infinitely many paths, and the limit must agree on all of them.

University

Partial derivatives

∂f∂xi\frac{\partial f}{\partial x_i}: the derivative with respect to one variable, holding the others fixed. Each answers "how sensitive is the output to this input?" — for a neural network, to this weight.

University

Gradient

∇f=(∂f∂x1,…,∂f∂xn)\nabla f = \left(\frac{\partial f}{\partial x_1}, \dots, \frac{\partial f}{\partial x_n}\right): the vector of all partial derivatives. It points in the direction of steepest ascent, its length is that steepest slope, and it is perpendicular to the level sets. Walk against it and you go downhill fastest.

University

Directional derivative

The rate of change of ff moving in a unit direction uu: Duf=∇f⋅uD_u f = \nabla f\cdot u. It proves that −∇f-\nabla f is the direction of steepest descent, and it is what forward-mode AD computes (a Jacobian–vector product).

University

Total differential and linearization

Near a point, a differentiable function is approximately linear: df=∑i∂f∂xidxi\dd f = \sum_i \frac{\partial f}{\partial x_i}\dd x_i. The graph has a tangent plane, and small input errors propagate linearly.

University

Multivariable chain rule

When a variable influences the output through several paths, add the contributions of every path, each the product of the local derivatives along it: ∂z∂x=∑i∂z∂ui∂ui∂x\frac{\partial z}{\partial x} = \sum_i \frac{\partial z}{\partial u_i}\frac{\partial u_i}{\partial x}. In matrix form, Jacobians multiply. This is exactly what backpropagation computes on a network's graph.

University

Jacobian matrix

For F:ℝn→ℝmF : \R^n \to \R^m, the m×nm \times n matrix of all partial derivatives ∂Fi/∂xj\partial F_i/\partial x_j: the best linear approximation of FF near a point. Its determinant measures how FF stretches volumes.

University

Hessian matrix

The matrix of second partial derivatives ∂2f/∂xi∂xj\partial^2 f/\partial x_i\partial x_j: the curvature of ff in every direction. Its eigenvalues classify critical points (minimum, maximum, saddle) and control how fast optimizers can go.

University

Extrema in several variables

Solve ∇f=0\nabla f = 0, then classify with the Hessian. For least squares ∥Xw−y∥2\norm{Xw - y}^2 this gives the normal equations X𝖳Xw=X𝖳yX^{\mathsf T}Xw = X^{\mathsf T}y — linear regression in closed form.

University

Multiple integrals and change of variables

Integrals over regions of ℝn\R^n: volumes, masses, probabilities of random vectors. Computed as iterated integrals (Fubini) and transformed with the Jacobian determinant: dx=∣det⁡J∣ du\dd x = |\det J|\,\dd u.

University

Curvature (basic differential geometry)

How fast a curve turns (κ=1/radius\kappa = 1/\text{radius} of the best-fitting circle) or how a surface bends (mean and Gaussian curvature). Graphics and CAD use it to judge smoothness, smooth meshes and design roads and rails.

Advanced

Where this area leads in computing

↑ ↓ to navigate · ↵ · Esc