control//state estimation//Kalman filter//Kalman gain

How much a filter lets a measurement move its estimate. The **Kalman gain** \(K\) is the fraction of the disagreement between sensor and prediction that the filter believes: how much it lets the measurement move the estimate and correct its assumption. The correction reads


How much a filter lets a measurement move its estimate. The Kalman gain KKK is the fraction of the disagreement between sensor and prediction that the filter believes: how much it lets the measurement move the estimate and correct its assumption. The correction reads

x^k=x^k−+K (zk−x^k−)=(1−K) x^k−+K zk.\hat x_k=\hat x_k^-+K\,(z_k-\hat x_k^-)=(1-K)\,\hat x_k^-+K\,z_k.x^k​=x^k−​+K(zk​−x^k−​)=(1−K)x^k−​+Kzk​.

Read aloud, new equals prediction plus confidence times disagreement. Rearranged, it is a weighted mean in which KKK is the sensor's weight and 1−K1-K1−K the model's, and the two always add up to one. K=0.9K=0.9K=0.9 believes the sensor strongly; K=0.1K=0.1K=0.1, not much.

The gain is not chosen by experiment. The filter computes it, which is why everything around it has a name: K=P−/(P−+R)K=P^-/(P^-+R)K=P−/(P−+R), where P−P^-P− is the doubt about the prediction and RRR the doubt about the sensor, both variances. If the prediction is a disaster and the sensor good (P−≫RP^-\gg RP−≫R), KKK tends to one and the filter keeps the sensor; in the opposite case it tends to zero and ignores it; equal doubts give one half. A prediction with P−=1P^-=1P−=1 against a sensor with R=100R=100R=100 gives K=1/101≈0.01K=1/101\approx0.01K=1/101≈0.01: the sensor barely moves the needle.

The model predicts 110 m, the GPS says 114 m, so the disagreement is 4 m. With K=0.25K=0.25K=0.25 the filter moves 1 m and lands on 111 m: the GPS thinks I am 4 m further ahead, I trust it somewhat but not completely. A superb GPS (K=0.9K=0.9K=0.9) takes it to 113.6 m, almost where the GPS says; an awful one (K=0.05K=0.05K=0.05) to 110.2 m, almost where the physics said. That is the entire philosophy.

Why not simply average? Averaging is K=0.5K=0.5K=0.5 forever. The filter changes the weights as its confidence changes, 0.95 prediction and 0.05 measurement at one moment, 0.2 and 0.8 later, because it keeps track of how much it knows about what it knows.

With many states and sensors the gain becomes K=P−HT(HP−HT+R)−1K=P^-H^{\mathsf T}(HP^-H^{\mathsf T}+R)^{-1}K=P−HT(HP−HT+R)−1. It looks ugly and means the same thing: how strongly should I correct myself using the sensor. P−HTP^-H^{\mathsf T}P−HT says how each state variable covaries with what the sensor sees, and the inverse divides by the total uncertainty of the disagreement; with H=1H=1H=1 it is P−/(P−+R)P^-/(P^-+R)P−/(P−+R) again (state-space model).

Written as Cov⁡(state,measurement) Var⁡(measurement)−1\operatorname{Cov}(\text{state},\text{measurement}),\operatorname{Var}(\text{measurement})^{-1}Cov(state,measurement)Var(measurement)−1, the gain is the coefficient of a linear regression: at every step the filter asks how far its state moves, on average, when the measurement moves one unit.

Two mediocre measurements beat a good one. A prediction of variance 16 corrected by a sensor of variance 4 leaves 3.2, below even the sensor's own 4, because information, the inverse of a variance, always adds, like two resistors in parallel (inverse-variance weighting).

?QuestionIs this just an exponential low-pass filter?

Almost. Exponential smoothing, x^←x^+α(z−x^)\hat x\leftarrow\hat x+\alpha(z-\hat x)x^←x^+α(z−x^), is literally the correction above with a fixed K=αK=\alphaK=α. A Kalman filter is an exponential low-pass whose α\alphaα is computed by itself, at every step, from the physics and the noise. In many systems KKK even settles to a constant (initial covariance), and then the steady-state filter is an exponential filter with the optimal α\alphaα. What the Kalman filter adds is the right α\alphaα, the right transient, many variables at once, and the uncertainty as an output.