mathematics//statistics//Mahalanobis distance

The Mahalanobis distance says how many standard deviations a point lies from a distribution, counting every uncertainty and every correlation at once, and it is what filters and anomaly detectors use to decide whether a reading is plausible. A GPS fix 15 m from where a filter expected it means nothing until it is known how sure the filter was: against an expected spread of 30 m it is ordinary, against 1 m it is absurd. The distance divides the disagreement by the spread before comparing it with anything.


The Mahalanobis distance says how many standard deviations a point lies from a distribution, counting every uncertainty and every correlation at once, and it is what filters and anomaly detectors use to decide whether a reading is plausible. A GPS fix 15 m from where a filter expected it means nothing until it is known how sure the filter was: against an expected spread of 30 m it is ordinary, against 1 m it is absurd. The distance divides the disagreement by the spread before comparing it with anything.

For one quantity it is the number of sigmas, squared: d2=(x−μ)2/σ2d^2=(x-\mu)^2/\sigma^2d2=(x−μ)2/σ2, so a reading 15 m off with σ=1\sigma=1σ=1 m has d2=225d^2=225d2=225. For several quantities the variance becomes the covariance matrix Σ\SigmaΣ, and its inverse does the dividing:

d2=(x−μ)T Σ−1 (x−μ).d^2=(x-\mu)^{\mathsf T}\,\Sigma^{-1}\,(x-\mu).d2=(x−μ)TΣ−1(x−μ).

Points with the same d2d^2d2 lie on an ellipse (an ellipsoid in more dimensions) shaped like the cloud of the data. The inverse of Σ\SigmaΣ stretches each direction by how little the data spread along it, so a small step against the correlation counts for more than a large step along it.

Euclidean distance gets correlated data wrong. Two sensors whose errors move together (two barometers seeing the same weather) rarely disagree, so a reading that is close in metres but breaks their usual agreement is far in Mahalanobis terms, and a reading that is far in metres but moves both together is near. Only the second measure knows which disagreements are normal (covariance).

The threshold comes from the chi-squared distribution. When the data are Gaussian and Σ\SigmaΣ is right, d2d^2d2 follows a χ2\chi^2χ2 distribution with as many degrees of freedom as there are quantities, mmm, with mean mmm. That fixes the cut-off for a chosen confidence: 9 for one quantity at three sigma (99.73 %), 9.21 for two quantities at 99 %. A point beyond it is either rare or not from this distribution.

The threshold is only as good as Σ\SigmaΣ. An underestimated covariance makes ordinary points look far and rejects good data; an overestimated one lets outliers through. With heavy-tailed data (a Cauchy distribution in the extreme) the chi-squared percentages no longer hold, though the ranking of points by distance is still useful.

In a running filter the point is the innovation and Σ\SigmaΣ its predicted covariance; the test that drops a reading beyond the threshold is innovation gating.