mathematics//statistics//Cauchy distribution

The Cauchy distribution is the standard example of noise that has no variance, and it marks where every method built on means and variances stops working. It looks like a bell from a distance, symmetric and peaked at its centre, but its tails fall so slowly that extreme values are not rare enough: the integral that defines the variance diverges, and even the mean does not exist. A filter, an average or a control chart that needs a \(\sigma\) has nothing to be given.


The Cauchy distribution is the standard example of noise that has no variance, and it marks where every method built on means and variances stops working. It looks like a bell from a distance, symmetric and peaked at its centre, but its tails fall so slowly that extreme values are not rare enough: the integral that defines the variance diverges, and even the mean does not exist. A filter, an average or a control chart that needs a σ\sigmaσ has nothing to be given.

It is not exotic. The ratio of two independent zero-mean Gaussians is Cauchy, and that ratio appears whenever a measurement is computed by dividing by something that can come close to zero: an angle from the quotient of two noisy components, a speed from a distance over a tiny time interval, a gain from a signal over a weak reference. Most of the time the denominator is comfortably large and the result is tame; once in a while it passes near zero and the result is enormous.

The practical symptom is a sample variance that never settles. With Gaussian data, the sample σ\sigmaσ of a few hundred readings is already close to the true one and only gets closer. With Cauchy data, it seems to settle for a while, then one monstrous value arrives and throws it upwards, and more data only means more such values.

after 100 drawsGaussian 0.97, Cauchy 19.9 after 1,000 drawsGaussian 1.03, Cauchy 21.8 after 20,000 drawsGaussian 1.00, Cauchy 127.8 The running sample standard deviation of a Gaussian with σ = 1 and of a Cauchy with scale 1, drawn one value at a time: the Gaussian settles on 1, the Cauchy jumps whenever a giant value arrives and never settles.

Play it and watch the copper line: each step up is one single draw, and a new run puts the steps elsewhere while the green line lands on 1 every time.

Averaging does not help. The mean of nnn Cauchy values is again Cauchy with the same width as one value, so a thousand readings average to something exactly as wild as a single one; this is the case where the central limit theorem does not apply, because its condition of finite variance fails.

Its width is still measurable, with statistics that ignore the tails. The median locates its centre and half the interquartile range equals its scale parameter, both of which settle normally; what fails is any statistic that squares the data.

For a Kalman filter it is the case that breaks everything, since the filter needs an RRR and there is none (Gaussian assumption). Real sensors are rarely exactly Cauchy, but a sensor that divides by a noisy quantity can come close enough that its sample variance keeps growing with the size of the recording, which is the sign to look for.