control//state estimation//Kalman filter//filter consistency//NIS

The NIS (normalized innovation squared) is the squared Mahalanobis distance of a filter's innovation, the surprise measured in units of the surprise the filter expected, and it is the standard consistency test of a running filter because it needs no ground truth, only the readings and what the filter predicted for them. Before the reading in blue, what the sensor brings in copper, the model in violet.


The NIS (normalized innovation squared) is the squared Mahalanobis distance of a filter's innovation, the surprise measured in units of the surprise the filter expected, and it is the standard consistency test of a running filter because it needs no ground truth, only the readings and what the filter predicted for them. Before the reading in blue, what the sensor brings in copper, the model in violet.

ϵk=ykTSk−1yk,Sk=HPk−HT+R.\epsilon_k=y_k^{\mathsf T}S_k^{-1}y_k,\qquad S_k=H\cb{P_k^-}H^{\mathsf T}+\cc{R}.ϵk​=ykT​Sk−1​yk​,Sk​=HPk−​HT+R.

yky_kyk​ is the innovation, the reading minus its prediction, and SkS_kSk​ how large it should be: the filter's doubt seen through the sensor plus the sensor's own noise. If the filter is consistent, ϵk\epsilon_kϵk​ follows a chi-square distribution with mmm degrees of freedom, mmm being the number of quantities measured at that step, so its mean is mmm. One value says little, since a chi-square with one degree of freedom is very spread; the test is the average over many steps. For one scalar measurement averaged over 100 steps, the mean NIS of an honest filter falls between 0.74 and 1.30 in 95 % of runs; a mean of 5 means real errors far larger than the filter believes, a mean of 0.2 a filter that expects surprises that never come.

Read the mean NIS against mmm. Ten times mmm means an arrogant filter, surprised far more than it expects, usually because Q\cd{Q}Q or R\cc{R}R is too small or the model is missing something. A tenth of mmm means a coward, expecting surprises that never come, with Q\cd{Q}Q or R\cc{R}R too large. Around mmm, sleep well.

One reading at a time, against a high threshold instead of averaged, the same number decides whether a measurement is believable at all (6.6 keeps 99 % of honest scalar readings), which is innovation gating.

It is how Q\cd{Q}Q and R\cc{R}R are tuned in practice: adjust them on real logs until the NIS behaves (filter tuning). A filter whose NIS grows only now and then is a different story: the world changed (a sensor broke, the vehicle took a load), and telling mistuning from change is the job of fault diagnosis.

It says that something is wrong, never what. A large NIS fits a sensor fault, a missing state, a delay and a wrong R\cc{R}R alike; the shape of the innovations over time narrows it down (filter consistency).