Partial derivatives

Level UniversityDifficulty ★★★★★Concept⌖ Open in the map

What is it?

∂f∂xi\frac{\partial f}{\partial x_i}: the derivative with respect to one variable, holding the others fixed. Each answers "how sensitive is the output to this input?" — for a neural network, to this weight.

Why does it exist?

With many inputs, "the" rate of change depends on the direction you move. The simplest directions are the axes: move one variable, freeze the rest. Partial derivatives are those axis-aligned rates, and (for nice functions) they determine the rate in every other direction.

Intuition

Slice the surface z=f(x,y)z = f(x, y) with the plane y=y0y = y_0: you get a curve, and ∂f/∂x\partial f/\partial x is its slope. Slice with x=x0x = x_0 for ∂f/∂y\partial f / \partial y. On a mountain: ∂f/∂x\partial f/\partial x is how steep it is walking due east, ∂f/∂y\partial f/\partial y due north.

Formal definition

∂f∂xi(a)=lim⁡h→0f(a+h ei)−f(a)h,\frac{\partial f}{\partial x_i}(a) = \lim_{h\to 0}\frac{f(a + h\,e_i) - f(a)}{h},

where eie_i is the ii-th unit vector. If all partials exist and are continuous near aa (f∈C1f \in C^1), ff is differentiable at aa; if f∈C2f \in C^2, mixed partials commute: ∂x∂yf=∂y∂xf\partial_x\partial_y f = \partial_y\partial_x f (Schwarz).

Formulas

∂f∂xi(a)=lim⁡h→0f(a+hei)−f(a)h\frac{\partial f}{\partial x_i}(a) = \lim_{h\to 0}\frac{f(a + h e_i) - f(a)}{h}
f(x,y)=x2y+sin⁡y:fx=2xy,fy=x2+cos⁡yf(x, y) = x^2 y + \sin y: \quad f_x = 2xy, \quad f_y = x^2 + \cos y

How is it computed?

Differentiate in xix_i treating every other variable as a constant; all single-variable rules apply.

Example

Squared error of a linear model with two weights: L(w,b)=(wx+b−y)2L(w, b) = (wx + b - y)^2. Then ∂L/∂w=2(wx+b−y) x\partial L/\partial w = 2(wx + b - y)\,x and ∂L/∂b=2(wx+b−y)\partial L/\partial b = 2(wx + b - y). The common factor 2(y^−y)2(\hat y - y) is the "error signal" that backpropagation sends to every weight.

Why does it matter?

Training computes ∂L/∂θi\partial L / \partial\theta_i for every parameter θi\theta_i; image processing computes ∂I/∂x\partial I/\partial x and ∂I/∂y\partial I/\partial y to find edges; physics writes its laws (heat, waves, fluids, electromagnetism, quantum mechanics) as equations between partial derivatives.

Where it shows up in computing

  • Image processing and computer vision★★★★★fundamentalSignals, media and vision

    Edge detection estimates ∂I/∂x\partial I/\partial x and ∂I/∂y\partial I/\partial y of the intensity.

  • Heat equation and diffusion★★★★★fundamentalPhysics and simulation

    The heat equation ut=α(uxx+uyy)u_t = \alpha(u_{xx} + u_{yy}) is a relation between partial derivatives.

Where it shows up in AI

  • Backpropagation★★★★★fundamentalAI and machine learning

    Backprop computes ∂L/∂w\partial L/\partial w for every weight ww of the network.

Where is it used?

Computing topics reachable from here, through the chain of ideas that leads to them:

What depends on it

Exercises

1Computation

Compute all first and second partial derivatives of f(x,y)=exyf(x, y) = e^{xy} and check that fxy=fyxf_{xy} = f_{yx}.

Solution

fx=yexyf_x = ye^{xy}, fy=xexyf_y = xe^{xy}, fxx=y2exyf_{xx} = y^2e^{xy}, fyy=x2exyf_{yy} = x^2e^{xy}, fxy=fyx=(1+xy)exyf_{xy} = f_{yx} = (1 + xy)e^{xy}.

2AI

For y^=σ(w1x1+w2x2+b)\hat y = \sigma(w_1x_1 + w_2x_2 + b) and L=−(yln⁡y^+(1−y)ln⁡(1−y^))L = -\big(y\ln\hat y + (1-y)\ln(1-\hat y)\big), show that ∂L/∂w1=(y^−y) x1\partial L/\partial w_1 = (\hat y - y)\,x_1.

Solution

∂L∂y^=y^−yy^(1−y^)\frac{\partial L}{\partial\hat y} = \frac{\hat y - y}{\hat y(1 - \hat y)}, ∂y^∂z=y^(1−y^)\frac{\partial\hat y}{\partial z} = \hat y(1 - \hat y), ∂z∂w1=x1\frac{\partial z}{\partial w_1} = x_1. The product simplifies to (y^−y)x1(\hat y - y)x_1: sigmoid and cross-entropy cancel beautifully.

↑ ↓ to navigate · ↵ · Esc