control//state estimation//Kalman filter//Gaussian assumption
Does a Kalman filter assume Gaussian noise or not? The answer has two levels. The equations do not need it: they use only means and covariances, the first two moments, and a variance is defined the same way for any distribution, \(\mathbb E[(X-\mu)^2]\). There is no Gaussian variance and non-Gaussian variance, which is why \(R\) is computed with the usual formula even for a sensor full of reflections. With any noise of finite variance, the filter is the **best linear unbiased estimator** (BLUE): the best of all estimators of the form prediction plus gain times disagreement.
Does a Kalman filter assume Gaussian noise or not? The answer has two levels. The equations do not need it: they use only means and covariances, the first two moments, and a variance is defined the same way for any distribution, E[(X−μ)2]\mathbb E[(X-\mu)^2]E[(X−μ)2]. There is no Gaussian variance and non-Gaussian variance, which is why RRR is computed with the usual formula even for a sensor full of reflections. With any noise of finite variance, the filter is the best linear unbiased estimator (BLUE): the best of all estimators of the form prediction plus gain times disagreement.
Gaussianity gives it two superpowers on top. It becomes optimal among all estimators, linear or not, because x^\hat xx^ is then exactly the mean of the Bayesian posterior and no algorithm can beat it in mean squared error (Bayes' rule). And PPP tells the whole story: the bell of (x^,P)(\hat x,P)(x^,P) is the complete belief, and ±2σ\pm2\sigma±2σ contains the truth 95 % of the time. Without Gaussianity PPP is still the right variance of the error, if the model is right, but no longer the whole story, and a nonlinear estimator could do better.
The exception that breaks everything is a distribution without variance. The Cauchy distribution, which appears as the ratio of two Gaussians (dividing by something that can approach zero), has tails so heavy that its variance is infinite: its sample variance never settles, because every so often a monstrous value arrives. There is no RRR to give the filter.
The fixes are of three kinds. Tweaks inside the method: gating the innovation against outliers, a robust update (Huber, Student-t) that shrinks large innovations instead of dropping them, an adaptive RRR estimated from the innovations, a state augmented with the bias (colored noise). Dirty tricks: inflating QQQ or RRR to make the filter humbler, a floor on PPP so it never stops listening. Macro methods: a bank of Kalman filters, one per hypothesis, weighted by probability (a Gaussian sum, or IMM for switching regimes), or a particle filter that represents the belief with thousands of samples of any shape, at a price.
Nonlinearity is a different problem with a different fix. Angles, rotations and distances to beacons are not Ax+bAx+bAx+b, and the extended and unscented filters (EKF, UKF) linearize with Jacobians or propagate sigma points. Almost everything real needs one of them.
More assumptions break in production: arithmetic in float32 on a microcontroller can make PPP lose symmetry, fixed by the Joseph form P=(I−KH)P−(I−KH)T+KRKTP=(I-KH)P^-(I-KH)^{\mathsf T}+KRK^{\mathsf T}P=(I−KH)P−(I−KH)T+KRKT; readings arrive late or out of order; nobody knows how correlated two sources are; and the cost grows with the cube of the state size, which is why weather forecasting uses an ensemble of samples instead.