What is it?
is the instantaneous rate of change of at : the slope of the tangent line to the graph, defined as the limit of slopes of secant lines.
Why does it exist?
Average rates are easy: distance over time. But "how fast is the car going now?" asks for a rate over an interval of length zero — . Newton needed it for mechanics, Leibniz for tangents; the derivative is their common answer, made rigorous by the limit. It solves the problem of quantifying local change.
Intuition
Geometrically: zoom in on a smooth curve and it looks like a straight line; the derivative is that line's slope. Physically: if is position, is velocity and acceleration. As sensitivity: — the output moves times as much as the input, for small nudges. That last reading is the one machine learning uses: the derivative of the loss with respect to a weight says which way to move the weight.
Formal definition
when the limit exists (then is differentiable at ). The function is the derivative of ; Leibniz writes it .
Formulas
- definition
- tangent line
- the basic table
How is it computed?
Three ways, and all three are used by software:
- Symbolically, applying rules (power, product, quotient, chain) to a formula — what a CAS does.
- Numerically, with a finite difference — easy but inexact (truncation error for large , rounding error for tiny ).
- Automatically: propagate exact derivatives through each elementary operation of a program — automatic differentiation, the engine of PyTorch and JAX.
Example
at : . So and the tangent is . In code, the central difference with gives , while gives garbage — the subtraction cancels all the digits.
Interactive visualization
Why does it matter?
Optimization is following derivatives downhill; simulation is integrating them forward in time; rendering needs them to orient surfaces; control uses them to anticipate. Modern deep learning is, at its computational core, an extremely efficient machine for computing derivatives of one number (the loss) with respect to billions of inputs (the weights).
Where it shows up in computing
Velocity is the derivative of position, acceleration the derivative of velocity.
Engines store positions and their derivatives (velocities) and advance them frame by frame.
The D term reacts to the derivative of the error, anticipating where the system is going.
Edges are where intensity changes fast: detectors (Sobel, Canny) estimate derivatives of the image.
Symbolic differentiation is the textbook example of term rewriting on expression trees.
Where it shows up in AI
Each step moves the parameter against the derivative of the loss: .
AD computes exact derivatives of programs by propagating them through each operation.
Where is it used?
Computing topics reachable from here, through the chain of ideas that leads to them:
λ Scientific computing and algorithms
- Differentiation rules→Symbolic computation (CAS)★★★★★
- Newton's method→Scientific computing★★★★★
- Differential and linear approximation→Conditioning→Numerical stability→Floating point (IEEE 754)★★★★★
- Antiderivatives and indefinite integrals→Fundamental theorem of calculus→Cumulative distribution function→Monte Carlo methods★★★★★
- Antiderivatives and indefinite integrals→Fundamental theorem of calculus→Improper integrals→Integral test→Algorithm analysis and complexity★★★★★
3D Computer graphics
- Partial derivatives→Gradient→Surface normals★★★★★
- Partial derivatives→Gradient→Signed distance fields and ray marching★★★★★
- Higher-order derivatives→Curvature (basic differential geometry)→Mesh processing (discrete differential geometry)★★★★★
- Partial derivatives→Gradient→Surface normals→Lighting and shading★★★★★
- Partial derivatives→Gradient→Surface normals→Ray tracing★★★★★
- Partial derivatives→Gradient→Surface normals→Ray tracing→The rendering equation★★★★★
- +1
⚙ Robotics and control
- Kinematics: position, velocity, acceleration★★★★★
- Kinematics: position, velocity, acceleration→Robot Jacobian (velocity kinematics)★★★★★
- Kinematics: position, velocity, acceleration→Robot dynamics★★★★★
- Ordinary differential equations→Control theory★★★★★
- Kinematics: position, velocity, acceleration→Robot Jacobian (velocity kinematics)→Inverse kinematics★★★★★
- Kinematics: position, velocity, acceleration→Robot dynamics→Trajectory optimization and MPC★★★★★
- +2
∿ Signals, media and vision
- Partial derivatives→Image processing and computer vision★★★★★
- Partial derivatives→Image processing and computer vision→Media compression (JPEG, MP3, video)★★★★★
- Ordinary differential equations→Laplace transform→Z-transform→Digital filters★★★★★
- Antiderivatives and indefinite integrals→Integration by substitution→Trigonometric integrals→Fourier series→Signal processing★★★★★
- Antiderivatives and indefinite integrals→Fundamental theorem of calculus→Improper integrals→Fourier transform→Sampling theorem (Nyquist–Shannon)★★★★★
- Antiderivatives and indefinite integrals→Fundamental theorem of calculus→Improper integrals→Fourier transform→Fast Fourier transform (FFT)★★★★★
- +1
ψ Quantum computing and physics
- Partial derivatives→Divergence→Laplacian→Schrödinger equation★★★★★
- Antiderivatives and indefinite integrals→Fundamental theorem of calculus→Improper integrals→Fourier transform→Uncertainty principle★★★★★
- Higher-order derivatives→Taylor polynomial→Taylor's theorem and the remainder→Taylor and Maclaurin series→Quantum computing★★★★★
What depends on it
- Differentiability and one-sided derivatives
- Differentiation rules
- Chain rule
- Higher-order derivatives
- Critical points and derivative tests
- Antiderivatives and indefinite integrals
- Differential and linear approximation
- Newton's method
- Numerical differentiation
- Partial derivatives
- Ordinary differential equations
Exercises
Using the definition, compute for .
Solution
.
Sketch and, below it, . Where is zero, positive, negative?
Solution
: zero at (a local max at , a local min at ), negative on where decreases, positive outside.
A one-parameter model has loss . Starting at with learning rate , compute two steps of gradient descent.
Solution
. ; . Each step closes 20% of the gap to the minimum .