mathematics//statistics//normal distribution//multivariate normal distribution
The multivariate normal distribution is the extension of the normal distribution to a vector of quantities, fixed entirely by a mean vector and a covariance matrix, and it is the belief every Kalman filter carries and the model of normal behaviour in most multivariable monitors: the position fix of a GNSS receiver, the joint readings of a motor's temperature and current. For a vector \(x\) of dimension \(k\) with mean \(\mu\) and covariance \(\Sigma\),
The multivariate normal distribution is the extension of the normal distribution to a vector of quantities, fixed entirely by a mean vector and a covariance matrix, and it is the belief every Kalman filter carries and the model of normal behaviour in most multivariable monitors: the position fix of a GNSS receiver, the joint readings of a motor's temperature and current. For a vector xxx of dimension kkk with mean μ\muμ and covariance Σ\SigmaΣ,
p(x) ∝ exp (−12 d2),d2=(x−μ)TΣ−1(x−μ).p(x)\;\propto\;\exp\!\left(-\tfrac12\,d^2\right),\qquad d^2=(x-\mu)^{\mathsf T}\Sigma^{-1}(x-\mu).p(x)∝exp(−21d2),d2=(x−μ)TΣ−1(x−μ).
ddd is the Mahalanobis distance, the distance from the mean counted in standard deviations along every direction at once. The density is the same wherever d2d^2d2 is the same, so its contours are ellipses (ellipsoids in more dimensions), centred on μ\muμ, with axes along the eigenvectors of Σ\SigmaΣ and half-lengths proportional to the square roots of its eigenvalues. A GNSS fix is such an ellipse, a few metres across, usually longer in the vertical than in the horizontal because the satellites are all above the receiver.
The tilt of the ellipse is what one variable at a time cannot see. A motor at 2σ2\sigma2σ in temperature and 2σ2\sigma2σ in current is ordinary if both rose together (the motor is working harder) and very rare if one rose while the other fell, though each reading alone looks the same.
It survives linear maps. If x∼N(μ,Σ)x\sim\mathcal N(\mu,\Sigma)x∼N(μ,Σ), then Ax+b∼N(Aμ+b, AΣAT)Ax+b\sim\mathcal N(A\mu+b,,A\Sigma A^{\mathsf T})Ax+b∼N(Aμ+b,AΣAT), and sums of independent Gaussian vectors are Gaussian. That closure is why a Kalman filter needs only a mean and a covariance per step, and why its prediction is Σ\SigmaΣ pushed through the dynamics (covariance propagation).
Conditioning keeps it Gaussian too. Knowing part of the vector leaves the rest Gaussian, with a mean moved in proportion to the surprise in the known part and a smaller covariance; with the state and a reading as the two parts, that conditioning is exactly the Kalman update (Kalman gain).
Zero correlation means independence here. For jointly Gaussian variables a zero off-diagonal entry implies independence, a step that is not valid in general for other distributions (covariance).
Its squared distance is χk2\chi^2_kχk2, which is where the thresholds of gates and monitors come from (chi-square distribution); its eigen-axes are the principal components of PCA. Like the one-dimensional bell, it is reasonable in the centre and optimistic in the tails.