computer science//floating-point arithmetic
Floating-point arithmetic is the way computers represent real numbers, as a fixed number of significant binary digits (the mantissa) times a power of two (the exponent), and it is what lets one 32-bit or 64-bit format hold both a bearing's micrometre of play and a grid's gigawatts with the same relative precision. The precision is relative and finite: a `float32` keeps about 7 significant decimal digits and a `float64` about 16, so every operation rounds its result to the nearest representable number, with a relative error of about \(10^{-7}\) or \(10^{-16}\) (the **machine epsilon**).
Floating-point arithmetic is the way computers represent real numbers, as a fixed number of significant binary digits (the mantissa) times a power of two (the exponent), and it is what lets one 32-bit or 64-bit format hold both a bearing's micrometre of play and a grid's gigawatts with the same relative precision. The precision is relative and finite: a float32 keeps about 7 significant decimal digits and a float64 about 16, so every operation rounds its result to the nearest representable number, with a relative error of about 10−710^{-7}10−7 or 10−1610^{-16}10−16 (the machine epsilon).
Rounding is invisible until two things happen. The first is adding very different magnitudes: in float32, adding 5×10−85\times10^{-8}5×10−8 to an accumulator that holds 1 changes nothing, because the increment is below half the spacing between representable numbers near 1. An integrator or a timestamp counter that adds small increments to a large running value in single precision silently stops moving. The second is catastrophic cancellation, subtracting two nearly equal numbers: the leading digits cancel and what is left is mostly the rounding of the operands. Computing a variance as E[x2]−E[x]2\mathbb E[x^2]-\mathbb E[x]^2E[x2]−E[x]2 for a barometer at 101,325 Pa with 1 Pa of noise subtracts two eleven-digit numbers to get a one-digit answer, and in float32 it comes out as garbage, sometimes negative (Welford's algorithm avoids it).
A problem's condition number spends the digits.
Solving a system loses about log10κ\log_{10}\kappalog10κ significant digits, so a condition number of a million leaves a float32 computation on a microcontroller with roughly one reliable digit, and a float64 one with ten.
Single precision is the embedded default. Microcontrollers in the Cortex-M4 and M7 class usually carry an optional hardware unit for float32 (some M7 parts for float64 too), and doubles elsewhere run in slow software. A Kalman filter in single precision can see its covariance lose symmetry or go negative after hours of updates; square-root and factorized forms exist for that reason. Neural network inference goes the other way, down to 16-bit and 8-bit formats, because it tolerates rounding well (quantization).
Equality tests are a trap. 0.1 + 0.2 == 0.3 is false in binary floating point, so code compares with a tolerance scaled to the magnitudes involved.
The order of operations changes the result, because rounding does not commute with reordering: a sum over a million samples accumulated naively drifts more than one summed in pairs, and a parallel reduction on a GPU can give slightly different answers from run to run.
Where there is no floating-point hardware, numbers are stored as scaled integers instead, and the programmer chooses the scale (fixed-point arithmetic). How rounding grows inside an algorithm is the subject of numerical methods.