control//state estimation//Kalman filter//correlated measurements

With several sensors, what a Kalman filter has to do is basically this: never count the same information as if it were two independent pieces of evidence. When the sensors are independent, their information simply adds, \(1/P=1/P^-+1/R_1+1/R_2\) (inverse-variance weighting), and the sum is commutative: stacking both readings in one vector \(z\) with a diagonal \(R\), or feeding them one after the other, gives exactly the same result. For independent sensors, batch equals sequential, and the order does not matter.


With several sensors, what a Kalman filter has to do is basically this: never count the same information as if it were two independent pieces of evidence. When the sensors are independent, their information simply adds, 1/P=1/P−+1/R1+1/R21/P=1/P^-+1/R_1+1/R_21/P=1/P−+1/R1​+1/R2​ (inverse-variance weighting), and the sum is commutative: stacking both readings in one vector zzz with a diagonal RRR, or feeding them one after the other, gives exactly the same result. For independent sensors, batch equals sequential, and the order does not matter.

The trap is plugging in the same sensor twice. Two readings identical to the decimal, and a filter that treats them as independent concludes: wow, how precise, two sensors of σ=2\sigma=2σ=2 m agree exactly, so my uncertainty drops to 1.41 m. False. It is one opinion with an echo, the real σ\sigmaσ is still 2 m, and the filter has become overconfident. The covariance C12C_{12}C12​ between the two errors cannot be ignored, so the readings go in as one vector with the full, square RRR:

z=[z1z2],R=[σ12ρ σ1σ2ρ σ1σ2σ22].z=\begin{bmatrix}z_1\\z_2\end{bmatrix},\qquad R=\begin{bmatrix}\sigma_1^2&\rho\,\sigma_1\sigma_2\\\rho\,\sigma_1\sigma_2&\sigma_2^2\end{bmatrix}.z=[z1​z2​​],R=[σ12​ρσ1​σ2​​ρσ1​σ2​σ22​​].

With σ1=2\sigma_1=2σ1​=2, σ2=3\sigma_2=3σ2​=3 and ρ=0.35\rho=0.35ρ=0.35 the covariance is 0.35⋅2⋅3=2.10.35\cdot2\cdot3=2.10.35⋅2⋅3=2.1, so RRR has 4 and 9 on the diagonal and 2.1 off it, ready to be used in the filter as RRR. Inside the gain K=P−HT(HP−HT+R)−1K=P^-H^{\mathsf T}(HP^-H^{\mathsf T}+R)^{-1}K=P−HT(HP−HT+R)−1, the inverse of that matrix accounts for the part of sensor 1 and sensor 2 that is redundant, and tells the filter how much of the information is really new.

weightsclaims σreal σ Ignoring ρ0.69 / 0.311.661.91 Full R0.78 / 0.221.891.89 Ignoring ρ: overconfident by 15 %. Two sensors with σ1 = 2, σ2 = 3 and error correlation ρ = 0.35.

With those numbers, fusing the two alone, the naive method weights them 0.69 and 0.31 and claims ±1.66\pm1.66±1.66 m while really erring ±1.91\pm1.91±1.91 m; the full RRR weights them 0.78 and 0.22 and claims and delivers ±1.89\pm1.89±1.89 m. One of the most important conclusions of the whole subject.

A negative correlation makes the naive filter underconfident instead, wasting information. And when the bad sensor is correlated enough with the good one (ρ>σ1/σ2\rho>\sigma_1/\sigma_2ρ>σ1​/σ2​), the optimal weight of the bad one turns negative: sharing almost the same error, it becomes a gauge of the common error, which is subtracted. Differential GPS and noise-cancelling headphones live on that idea.

Two GPS receivers on one car share ionosphere, satellites and reflections, so a second one does not buy the 2\sqrt22​. Two sensors on one power supply share its ripple. A GPS and a wheel odometer rest on different physics and err almost independently. Engineering prefers diverse sensors to duplicated ones for this reason.

The same mistake in time, one sensor whose successive errors resemble each other, is colored noise.

?QuestionAnd if I still want to feed them one at a time?

You can, after whitening them. Factor R=LLTR=LL^{\mathsf T}R=LLT (Cholesky), transform z′=L−1zz'=L^{-1}zz′=L−1z and H′=L−1HH'=L^{-1}HH′=L−1H, and the new R′R'R′ is the identity: the transformed readings are independent, so the order is free again and sequential equals batch. What sequential processing cannot do is skip that step, because fed raw, one after the other, the second reading is counted as fresh evidence when part of it is the first one's echo.