mathematics//signal processing//quantization

Quantization is the rounding of each sample of a signal to the nearest of a finite set of levels, and it is the second cut every digital measurement suffers after sampling: an analog-to-digital converter with \(N\) bits can only say which of \(2^N\) levels a voltage is closest to. The same rounding reappears wherever numbers are stored with few bits, from a position sent over a link with 1 cm resolution to the weights of a neural network stored as 8-bit integers (model quantization, the same mathematics applied to stored numbers instead of a physical input).


Quantization is the rounding of each sample of a signal to the nearest of a finite set of levels, and it is the second cut every digital measurement suffers after sampling: an analog-to-digital converter with NNN bits can only say which of 2N2^N2N levels a voltage is closest to. The same rounding reappears wherever numbers are stored with few bits, from a position sent over a link with 1 cm resolution to the weights of a neural network stored as 8-bit integers (model quantization, the same mathematics applied to stored numbers instead of a physical input).

The size of one level, the least significant bit or step Δ\DeltaΔ, is the range divided among the levels, and when the signal moves enough between samples the rounding error behaves like noise spread evenly between −Δ/2-\Delta/2−Δ/2 and +Δ/2+\Delta/2+Δ/2:

Δ=VFS2N,σq=Δ12.\Delta=\frac{V_{FS}}{2^N},\qquad \sigma_q=\frac{\Delta}{\sqrt{12}} .Δ=2NVFS​​,σq​=12​Δ​.

VFSV_{FS}VFS​ is the full-scale range and σq\sigma_qσq​ the standard deviation of the quantization noise. A 12-bit converter over 3.3 V has steps of 0.81 mV and a quantization noise of 0.23 mV. A 16-bit gyro set to ±2000\pm 2000±2000 °/s resolves 0.061 °/s per count, and a 16-bit accelerometer at ±16\pm 16±16 g resolves 0.49 mg.

Range and resolution are bought from the same bits.

A wider range survives the shocks of a hard landing without saturating, and makes every step coarser; a narrow range resolves finely and clips. A signal that never uses more than a tenth of the range has thrown away more than three bits before any noise is counted.

The uniform-noise model needs motion. A slow or constant signal sits in one step and the error becomes a fixed offset that depends on the value, correlated from sample to sample; averaging then does not shrink it. Adding a little random noise before the converter (dither) restores the uniform behaviour, which is why a sensor with some analog noise can be averaged below one step and a perfectly clean one cannot (white noise).

Each bit adds about 6 dB to the best possible signal-to-noise ratio, because halving Δ\DeltaΔ halves σq\sigma_qσq​. Past the noise of the analog stage in front, further bits only write that noise with more digits (signal-to-noise ratio).

In an error budget it is usually the smallest term. A position quantized to 1 cm contributes 0.01/12≈0.0030.01/\sqrt{12}\approx 0.0030.01/12​≈0.003 m, beside a latency term that can be a hundred times larger, which is why buying bits rarely fixes a system (error budget).

Resolution is the step, accuracy is how close the reading is to the truth; a fine step says nothing about gain error or bias (precision and accuracy).