Taylor polynomial

Level UniversityDifficulty ★★★★★Concept⌖ Open in the map

What is it?

The polynomial of degree nn that matches ff and its first nn derivatives at a point aa. Degree 1 is the tangent line, degree 2 adds curvature; the higher the degree, the wider the region where it is a good approximation. At a=0a = 0 it is called a Maclaurin polynomial.

Why does it exist?

Polynomials are what we can compute: additions and multiplications. Other functions — sin⁡\sin, exe^x, the solution of a differential equation — are not. Taylor's idea gives a systematic way to trade any smooth function for a polynomial near a point, with a precise error estimate.

Intuition

Matching more derivatives means matching the shape more closely: value (same height), first derivative (same slope), second (same bending), third (same change of bending)… Each extra term corrects what the previous polynomial got wrong, and the corrections get smaller near aa because they carry higher powers of (x−a)(x - a). Move the slider in the demo and watch the approximation "peel" away from aa more slowly as the order grows.

Formal definition

If ff is nn times differentiable at aa, its Taylor polynomial of order nn is

Tn(x)=∑k=0nf(k)(a)k!(x−a)k,T_n(x) = \sum_{k=0}^{n}\frac{f^{(k)}(a)}{k!}(x - a)^k,

the unique polynomial of degree ≤n\le n with Tn(k)(a)=f(k)(a)T_n^{(k)}(a) = f^{(k)}(a) for k=0,…,nk = 0, \dots, n. Moreover f(x)=Tn(x)+o((x−a)n)f(x) = T_n(x) + o\big((x - a)^n\big) as x→ax \to a.

Formulas

Tn(x)=f(a)+f′(a)(x−a)+f′′(a)2!(x−a)2+⋯+f(n)(a)n!(x−a)nT_n(x) = f(a) + f'(a)(x-a) + \frac{f''(a)}{2!}(x-a)^2 + \dots + \frac{f^{(n)}(a)}{n!}(x-a)^n
ex≈1+x+x22+x36,sin⁡x≈x−x36+x5120e^x \approx 1 + x + \frac{x^2}{2} + \frac{x^3}{6}, \qquad \sin x \approx x - \frac{x^3}{6} + \frac{x^5}{120}
f(x+h)≈f(x)+∇f(x)⋅h+12 h𝖳∇2f(x) hf(x + h) \approx f(x) + \nabla f(x)\cdot h + \tfrac12\,h^{\mathsf T}\nabla^2 f(x)\,h
second order in several variables: gradient and Hessian

How is it computed?

Compute f(a),f′(a),f′′(a),…f(a), f'(a), f''(a), \dots, divide the kk-th by k!k!. For compositions it is usually faster to combine known series (substitute, multiply, integrate term by term) than to differentiate repeatedly. Software can do it exactly with Taylor-mode arithmetic on truncated power series — which is how the demo computes the coefficients.

Example

Small-angle approximation: sin⁡θ≈θ\sin\theta \approx \theta (order 1) turns the pendulum equation θ¨=−gLsin⁡θ\ddot\theta = -\frac gL\sin\theta into θ¨=−gLθ\ddot\theta = -\frac gL\theta, solvable by hand: period 2πL/g2\pi\sqrt{L/g}. The error at θ=10°≈0.1745\theta = 10° \approx 0.1745 is θ3/6≈0.0009\theta^3/6 \approx 0.0009 — about 0.5%.

Interactive visualization

f Tn— The shaded band is |x − a| < R. Outside it the polynomials eventually diverge however high the order — for 1/(1 + x²) because of the complex poles at ±i, invisible on the real line.

Why does it matter?

Gradient descent is "trust the order-1 Taylor model a little"; Newton's method is "jump to the minimum of the order-2 model". Numerical differentiation, ODE solvers (Euler, Runge–Kutta) and their error analysis are derived from Taylor expansions. Physics linearizes with them, and math libraries start from them.

Where it shows up in computing

  • Floating point (IEEE 754)★★★★★frequentScientific computing and algorithms

    Math libraries evaluate sin⁡\sin, exp⁡\exp, log⁡\log with polynomial approximations after range reduction (minimax, refined from Taylor).

  • Physics engines★★★★★frequentPhysics and simulation

    Integrators (Euler, Verlet) are truncated Taylor expansions of the motion in the time step.

Where it shows up in AI

  • Second-order (Hessian-based) optimization★★★★★fundamentalAI and machine learning

    Newton and trust-region methods minimize the quadratic Taylor model f+g𝖳h+12h𝖳Hhf + g^{\mathsf T}h + \frac12 h^{\mathsf T}Hh.

  • Gradient descent★★★★★fundamentalAI and machine learning

    The first-order model f(x−ηg)≈f(x)−η∥g∥2f(x - \eta g) \approx f(x) - \eta\norm g^2 is why a small step downhill decreases ff.

Where is it used?

Computing topics reachable from here, through the chain of ideas that leads to them:

What depends on it

Exercises

1Computation

Find the Maclaurin polynomial of order 4 of cos⁡x\cos x and use it to estimate cos⁡0.5\cos 0.5.

Solution

T4=1−x22+x424T_4 = 1 - \frac{x^2}{2} + \frac{x^4}{24}; T4(0.5)=0.87760416‾T_4(0.5) = 0.8776041\overline{6} vs cos⁡0.5=0.8775826\cos 0.5 = 0.8775826 (error 2⋅10−52 \cdot 10^{-5}).

2AI

Use the second-order Taylor model of L(w)L(w) to derive the step that minimizes it. Which method is this?

Solution

L(w+h)≈L+L′h+12L′′h2L(w + h) \approx L + L'h + \frac12 L''h^2; setting the derivative in hh to zero gives h=−L′/L′′h = -L'/L''. That is Newton's method for optimization.

↑ ↓ to navigate · ↵ · Esc