computer science//fixed-point arithmetic
Fixed-point arithmetic is a way of representing fractional numbers as plain integers with an implied scale factor chosen by the programmer, and it is how control and signal-processing code runs on microcontrollers, DSPs and FPGAs that have no floating-point unit, or where integer operations are cheaper and more predictable. A motor current between -16 A and +16 A can be stored in a 16-bit integer as milliamps times two, or more usually in a **Q format**: in Q15 a 16-bit integer stands for a number between -1 and just under 1, the integer divided by \(2^{15}\), with a resolution of about \(3\times10^{-5}\).
Fixed-point arithmetic is a way of representing fractional numbers as plain integers with an implied scale factor chosen by the programmer, and it is how control and signal-processing code runs on microcontrollers, DSPs and FPGAs that have no floating-point unit, or where integer operations are cheaper and more predictable. A motor current between -16 A and +16 A can be stored in a 16-bit integer as milliamps times two, or more usually in a Q format: in Q15 a 16-bit integer stands for a number between -1 and just under 1, the integer divided by 2152^{15}215, with a resolution of about 3×10−53\times10^{-5}3×10−5.
The scale is fixed per variable, which is the defining difference from floating-point arithmetic, where the exponent travels with each number. That makes the arithmetic exact, fast and deterministic (an addition is one integer addition, with no rounding as long as nothing overflows), and it moves the work to design time. Each variable needs a range large enough never to overflow and a resolution fine enough not to drown its signal, and those two pull against each other in a fixed number of bits.
In a microcontroller without floating point, scaling becomes part of the design.
Every gain, state and intermediate product of a PID or a filter gets its own format, chosen from the physical range of the quantity, and a change of sensor or gain can force the formats to be redone.
Multiplication is where the care goes. Two Q15 numbers multiply into a 32-bit Q30 result that must be shifted back, and the intermediate needs the wider register or it overflows; a filter coefficient close to 1 (a slow low-pass) needs many bits of resolution or the filter stops responding to small changes. Saturating arithmetic, which clamps at the maximum instead of wrapping to a large negative number, is standard in DSP instruction sets, and it matters in a loop because a wrapped overflow inside a control loop turns full positive thrust into full negative.
The rounding behaves like quantization. Every truncation adds a small uniform error, the same s2/12s^2/12s2/12 variance as an ADC step of size sss (quantization), and a recursive filter feeds that error back, where it can sustain small limit cycles that a floating-point version would not show.
It is less common in new designs than it was. Microcontrollers with single-precision hardware (Cortex-M4F and up) are cheap, and libraries such as CMSIS-DSP offer both versions of each routine. Fixed point remains standard in the cheapest MCUs, in FPGAs, in motor drives running current loops at tens of kilohertz, and in 8-bit integer neural network inference, which is the same idea applied to weights (quantization).