What is it?
The finite approximation of the real numbers used by every computer: with a 53-bit mantissa in double precision. About 16 significant digits, relative rounding error at most , and a zoo of pitfalls: cancellation, overflow, non-associativity.
Why does it exist?
Calculus lives in , where limits exist and can tend to 0. A computer has finitely many numbers, spaced further apart as they grow. IEEE 754 (1985, Kahan) defines that set and guarantees each operation is correctly rounded, which is what makes numerical analysis — and reproducible science — possible.
Formulas
- floating-point addition is not associative
Why does it matter?
Every numerical topic in this portal meets floating point eventually: derivatives by finite differences lose digits, long sums drift, chaotic simulations diverge between machines, and machine learning trains in 16-bit and even 8-bit formats where these issues are front and centre.
The mathematics behind it
IEEE 754 doubles are a finite, unevenly spaced subset of ; every operation rounds back into it.
Floating point is designed around a guaranteed relative error per operation.
Catastrophic cancellation is the classic floating-point pitfall; stable reformulations avoid it.
Multiplying many small probabilities underflows to 0; summing their logs (log-sum-exp) does not.
Hardware and libraries compute and from a table guess refined by Newton iterations.
Math libraries evaluate , , with polynomial approximations after range reduction (minimax, refined from Taylor).
Outside the domain IEEE 754 returns NaN (
sqrt(-1)) or ±∞ (log(0)) instead of failing.Approaching a limit with tiny steps in floating point triggers catastrophic cancellation.
expm1(x)andlog1p(x)exist because computing or directly loses all precision for tiny .Library designers bound the remainder to decide how many terms reach full double precision.
Float addition is not associative; parallel reductions reorder terms and give run-to-run differences — order matters.
Rounding differences (operation order, FMA, GPU) are amplified exponentially in chaotic simulations.
IEEE 754 has :
1.0/0.0isinf, a finite stand-in for an infinite limit.Summing alternating terms of large size cancels digits; evaluating by its series in floating point gives garbage.
Where is it used?
Computing topics reachable from here, through the chain of ideas that leads to them:
What depends on it
Exercises
Why does evaluating for with get worse when goes below about ?
Solution
The truncation error is but the rounding error of the numerator is . Their sum is minimized at ; smaller amplifies rounding.
This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.