control//state estimation//Kalman filter//information filter
The information filter is the Kalman filter rewritten in terms of the inverse of the covariance, the **information matrix** \(Y=P^{-1}\) (how much is known, where the covariance says how much is doubted), and of the information vector \(\eta=P^{-1}\hat x\), so that the contributions of independent sensors simply add. It is the form of choice when many sensors report at once, when the fusion is spread over several computers, and when a filter must start knowing nothing at all.
The information filter is the Kalman filter rewritten in terms of the inverse of the covariance, the information matrix Y=P−1Y=P^{-1}Y=P−1 (how much is known, where the covariance says how much is doubted), and of the information vector η=P−1x^\eta=P^{-1}\hat xη=P−1x^, so that the contributions of independent sensors simply add. It is the form of choice when many sensors report at once, when the fusion is spread over several computers, and when a filter must start knowing nothing at all.
For independent measurements zi=Hix+viz_i=H_ix+v_izi=Hix+vi, each with noise covariance RiR_iRi, the correction is a sum:
Y=Y0+∑i=1NHiTRi−1Hi,η=η0+∑i=1NHiTRi−1zi,x^=Y−1η.Y=Y_0+\sum_{i=1}^{N}H_i^{\mathsf T}R_i^{-1}H_i,\qquad \eta=\eta_0+\sum_{i=1}^{N}H_i^{\mathsf T}R_i^{-1}z_i,\qquad \hat x=Y^{-1}\eta .Y=Y0+i=1∑NHiTRi−1Hi,η=η0+i=1∑NHiTRi−1zi,x^=Y−1η.
Y0Y_0Y0 and η0\eta_0η0 are what was known before, each term in the sums is what sensor iii brings, and the estimate is recovered by one solve at the end. It is inverse-variance weighting in matrix form: in one dimension, informations add like the conductances of resistors in parallel, and the fused variance is smaller than the smallest.
The two forms trade costs. Correcting is cheap in information form, an addition per sensor, and predicting is expensive, because pushing the information through the dynamics needs an inversion; the covariance form is the other way round. With many sensors and few predictions the information form wins.
Knowing nothing is exact here. A diffuse start, infinite covariance, is simply Y0=0Y_0=0Y0=0, which the covariance form can only approximate with a very large number (initial covariance).
The sum is what makes distributed fusion possible: a sum is NNN times an average, and a network of agents can compute an average among neighbours (distributed Kalman filter). In large SLAM problems the information matrix is sparse, since each observation links only a few poses and landmarks, while the covariance is dense; graph-based solvers exploit that sparsity.
The addition is only valid for independent errors. Two contributions that share an error, or one that already contains the other, must not be summed (correlated measurements).