mathematics//statistics//inverse-variance weighting

How two imperfect estimates of the same quantity combine into one better than both. If the estimates are independent, with variances \(P^-\) and \(R\), the best linear combination weights each by the inverse of its variance, and its variance is


How two imperfect estimates of the same quantity combine into one better than both. If the estimates are independent, with variances P−P^-P− and RRR, the best linear combination weights each by the inverse of its variance, and its variance is

P=P−RP−+R,1P=1P−+1R.P=\frac{P^-R}{P^-+R},\qquad \frac1P=\frac1{P^-}+\frac1R.P=P−+RP−R​,P1​=P−1​+R1​.

Intuition says the result should land between the two uncertainties, as if the fusion only averaged them. It does not. A prediction of variance 16 and a sensor of variance 4 combine into variance 3.2 (standard deviations 4 and 2 give 1.79), smaller even than 4. Two independent sources of information beat either one alone, even when one of them is fairly bad. The second form says why: the inverse of a variance is information, or precision in Bayesian language, and information adds, always, so the combined variance always sits below the smaller of the two.

K0.80 x̂13.20 P3.20 (σ 1.79) 1/P0.313 = 0.063 + 0.250 Prediction N(10, 16) in blue, measurement N(14, 4) in copper, fused estimate in green: narrower than either.

The green bell is always the narrowest. Drag the sensor further away and it moves, but it does not widen.

In a Kalman filter this is the covariance update in disguise: (1−K)P−(1-K)P^-(1−K)P− with K=P−/(P−+R)K=P^-/(P^-+R)K=P−/(P−+R) is exactly P−R/(P−+R)P^-R/(P^-+R)P−R/(P−+R) (estimate covariance).

The measured value zzz appears nowhere in the formula. In a linear filter the new variance depends on how noisy the sensor is, never on what it said, so a reading ten standard deviations away makes the filter exactly as much more confident as a sensible one. It is never surprised; the defence is gating the innovation.

It is the formula of two resistors in parallel, Req=R1R2/(R1+R2)R_{eq}=R_1R_2/(R_1+R_2)Req​=R1​R2​/(R1​+R2​), and for the same reason the equivalent is below the smaller one: two paths for the current, two paths for the information.

Independence is the price. Two readings that share an error are not two sources, and adding their information makes the result overconfident (correlated measurements).