control//state estimation//Kalman filter//robust Kalman filter
A robust Kalman filter keeps a sensor with heavy-tailed noise from dragging the estimate around: instead of trusting every reading in proportion to \(R\), it trusts large disagreements less than small ones. It is the tool for noise that is mostly a bell but produces big errors far more often than a bell would (a GPS in a city, a range sensor that sometimes catches the wrong surface, an inertial unit on a vibrating frame), where a Kalman filter is pulled by each large error as if it were information.
A robust Kalman filter keeps a sensor with heavy-tailed noise from dragging the estimate around: instead of trusting every reading in proportion to RRR, it trusts large disagreements less than small ones. It is the tool for noise that is mostly a bell but produces big errors far more often than a bell would (a GPS in a city, a range sensor that sometimes catches the wrong surface, an inertial unit on a vibrating frame), where a Kalman filter is pulled by each large error as if it were information.
The ordinary update moves the estimate by a fixed fraction of the innovation, however large: a reading ten sigmas away moves it ten times as far as a reading one sigma away. A robust update bends that line. Small innovations are used as usual; beyond a few sigmas, the reading's weight falls, which is the same as inflating its RRR for that step only. A reading twenty sigmas away is not thrown out, it is heard quietly.
The two usual recipes come from robust statistics. Huber's M-estimation treats the error as squared near zero and as an absolute value beyond a threshold, so the influence of a large error stops growing. A Student-t model of the noise does the same through its heavier tails, and its degrees of freedom set how heavy; with few degrees of freedom it ignores outliers almost completely, with many it becomes the Kalman filter again.
The price is that the update is no longer one closed line: the weight depends on the innovation, which depends on the estimate, so the step is solved by a short iteration or an approximation (variational methods for the Student-t), and PPP no longer comes out exact.
It differs from innovation gating in kind. Gating is all or nothing and suits rare, discrete faults; down-weighting suits a continuous heavy tail, where there is no clean line between an outlier and an unlucky reading.
It is a fix for the shape of the noise, not for its size. A sensor whose precision changes with the conditions needs a varying RRR (adaptive Kalman filter); a sensor that switches between two behaviours needs a filter per behaviour (IMM).