Floating point (IEEE 754)

Level FundamentalDifficulty ★★★★★Application⌖ Open in the map

What is it?

The finite approximation of the real numbers used by every computer: ±m×2e\pm m \times 2^e with a 53-bit mantissa in double precision. About 16 significant digits, relative rounding error at most 2−532^{-53}, and a zoo of pitfalls: cancellation, overflow, non-associativity.

Why does it exist?

Calculus lives in ℝ\R, where limits exist and hh can tend to 0. A computer has finitely many numbers, spaced further apart as they grow. IEEE 754 (1985, Kahan) defines that set and guarantees each operation is correctly rounded, which is what makes numerical analysis — and reproducible science — possible.

Formulas

x=(−1)s×1.f×2e−1023,εmach=2−52≈2.2×10−16x = (-1)^s \times 1.f \times 2^{e - 1023}, \qquad \varepsilon_{\text{mach}} = 2^{-52} \approx 2.2\times10^{-16}
(1017+1)−1017=0≠1=(1017−1017)+1(10^{17} + 1) - 10^{17} = 0 \ne 1 = (10^{17} - 10^{17}) + 1
floating-point addition is not associative

Why does it matter?

Every numerical topic in this portal meets floating point eventually: derivatives by finite differences lose digits, long sums drift, chaotic simulations diverge between machines, and machine learning trains in 16-bit and even 8-bit formats where these issues are front and centre.

The mathematics behind it

  • Real numbers★★★★★fundamental

    IEEE 754 doubles are a finite, unevenly spaced subset of ℝ\R; every operation rounds back into it.

  • Absolute and relative error★★★★★fundamental

    Floating point is designed around a guaranteed relative error per operation.

  • Numerical stability★★★★★fundamental

    Catastrophic cancellation is the classic floating-point pitfall; stable reformulations avoid it.

  • Logarithmic functions★★★★★frequent

    Multiplying many small probabilities underflows to 0; summing their logs (log-sum-exp) does not.

  • Newton's method★★★★★frequent

    Hardware and libraries compute 1/x1/x and x\sqrt x from a table guess refined by Newton iterations.

  • Taylor polynomial★★★★★frequent

    Math libraries evaluate sin⁡\sin, exp⁡\exp, log⁡\log with polynomial approximations after range reduction (minimax, refined from Taylor).

  • Domain and range★★★★★frequent

    Outside the domain IEEE 754 returns NaN (sqrt(-1)) or ±∞ (log(0)) instead of failing.

  • Limit of a function★★★★★frequent

    Approaching a limit with tiny steps in floating point triggers catastrophic cancellation.

  • Infinitesimals and equivalences★★★★★frequent

    expm1(x) and log1p(x) exist because computing ex−1e^x - 1 or ln⁡(1+x)\ln(1 + x) directly loses all precision for tiny xx.

  • Taylor's theorem and the remainder★★★★★frequent

    Library designers bound the remainder to decide how many terms reach full double precision.

  • Float addition is not associative; parallel reductions reorder terms and give run-to-run differences — order matters.

  • Rounding differences (operation order, FMA, GPU) are amplified exponentially in chaotic simulations.

  • Infinite limits★★★★★frequent

    IEEE 754 has ±∞\pm\infty: 1.0/0.0 is inf, a finite stand-in for an infinite limit.

  • Alternating series★★★★★frequent

    Summing alternating terms of large size cancels digits; evaluating e−20e^{-20} by its series in floating point gives garbage.

Where is it used?

Computing topics reachable from here, through the chain of ideas that leads to them:

What depends on it

Exercises

1Computing

Why does evaluating f′(1)f'(1) for f=exf = e^x with (f(1+h)−f(1))/h(f(1+h) - f(1))/h get worse when hh goes below about 10−810^{-8}?

Solution

The truncation error is ≈h2f′′\approx \frac{h}{2}f'' but the rounding error of the numerator is ≈εf/h\approx \varepsilon f/h. Their sum is minimized at h≈ε≈10−8h \approx \sqrt\varepsilon \approx 10^{-8}; smaller hh amplifies rounding.

This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.

↑ ↓ to navigate · ↵ · Esc