Lab

Lab

All the interactive visualizations in one place. Each one also lives on the page of its topic.

Derivative

FundamentalOpen the topic →

f′(a)f'(a) is the instantaneous rate of change of ff at aa: the slope of the tangent line to the graph, defined as the limit of slopes of secant lines.

f tangent, slope f′(a) secant through a and a + h— Drag the point. As h → 0 the secant becomes the tangent: that limit is the derivative.

Riemann sums

FundamentalOpen the topic →

Approximate the area under a curve by nn thin rectangles, ∑f(xi∗) Δx\sum f(x_i^\ast)\,\Delta x. As n→∞n \to \infty the sum converges to the integral — slowly for the left/right rule (O(1/n)O(1/n)), faster for the midpoint (O(1/n2)O(1/n^2)).

Increase n and watch the sum converge to the area. Left and right sums have error ∝ 1/n; midpoint and trapezoid ∝ 1/n².

Newton's method

UniversityOpen the topic →

To solve f(x)=0f(x) = 0, replace ff by its tangent line at the current guess and jump to where the tangent hits zero: xk+1=xk−f(xk)/f′(xk)x_{k+1} = x_k - f(x_k)/f'(x_k). Near a simple root the number of correct digits doubles every step.

kxkf(xk)|xk − x*|digits
Click the canvas to choose x₀. Near a simple root the number of correct digits roughly doubles each step. Try x³ − 2x + 2 from x₀ = 0 (a 2-cycle), ∛x (the iterates double and flee) or arctan x from |x₀| > 1.4.

Taylor polynomial

UniversityOpen the topic →

The polynomial of degree nn that matches ff and its first nn derivatives at a point aa. Degree 1 is the tangent line, degree 2 adds curvature; the higher the degree, the wider the region where it is a good approximation. At a=0a = 0 it is called a Maclaurin polynomial.

f Tn— The shaded band is |x − a| < R. Outside it the polynomials eventually diverge however high the order — for 1/(1 + x²) because of the complex poles at ±i, invisible on the real line.

Gradient descent

UniversityOpen the topic →

Repeat θ←θ−η ∇L(θ)\theta \leftarrow \theta - \eta\,\nabla L(\theta): take a small step against the gradient. Cauchy proposed it in 1847; today it (and its stochastic, adaptive variants) trains essentially every neural network.

Click anywhere to drop the ball there. Colour is height (log scale), lines are level curves; the gradient is perpendicular to them. Try the narrow valley with plain gradient descent, then with momentum.

Backpropagation

AdvancedOpen the topic →

The algorithm that computes ∂L/∂W\partial L/\partial W and ∂L/∂b\partial L/\partial b for every layer of a network: one forward pass storing intermediate values, then one backward pass applying the chain rule from the loss down to the inputs. Cost: about twice the forward pass, whatever the number of parameters.

x
w
b
→
z = wx + b
→
ŷ = f(z)
→
L = ½(ŷ − y)²∂L/∂L = 1

forward value ·gradient ∂L/∂· flowing backwards

The chain rule, evaluated backwards: ∂L/∂w = (∂L/∂ŷ)·(∂ŷ/∂z)·(∂z/∂w). For a layer ŷ = f(Wx + b) the same computation gives ∂L/∂W = δ xᵀ and ∂L/∂b = δ with δ = (ŷ − y) ⊙ f′(z). Try ReLU with a negative z: the gradient is zero and the neuron cannot learn ("dying ReLU").
↑ ↓ to navigate · ↵ · Esc