mathematics//probability//Bayes' rule

How a belief should change when evidence arrives. The belief before the evidence is the **prior** \(p(x)\); what the evidence says about each possible \(x\) is the **likelihood** \(p(z\mid x)\); the belief after is the **posterior**,


How a belief should change when evidence arrives. The belief before the evidence is the prior p(x)p(x)p(x); what the evidence says about each possible xxx is the likelihood p(z∣x)p(z\mid x)p(z∣x); the belief after is the posterior,

p(x∣z)  ∝  p(z∣x) p(x).p(x\mid z)\;\propto\;p(z\mid x)\,p(x).p(x∣z)∝p(z∣x)p(x).

The posterior is the prior reweighted by how well each candidate explains the evidence, then renormalized.

With Gaussians it becomes arithmetic. Believe x∼N(10,4)x\sim\mathcal N(10,4)x∼N(10,4) before measuring; a sensor with Gaussian noise of variance R=1R=1R=1 says z=12z=12z=12. The product of two Gaussians is a Gaussian, and its mean and variance come out of the Kalman filter equations: K=4/(4+1)=0.8K=4/(4+1)=0.8K=4/(4+1)=0.8, x^=10+0.8 (12−10)=11.6\hat x=10+0.8,(12-10)=11.6x^=10+0.8(12−10)=11.6, P=(1−0.8)⋅4=0.8P=(1-0.8)\cdot4=0.8P=(1−0.8)⋅4=0.8. The belief moves from N(10,4)\mathcal N(10,4)N(10,4) to N(11.6,0.8)\mathcal N(11.6,0.8)N(11.6,0.8): towards the sensor, and narrower (inverse-variance weighting).

A Kalman filter is recursive Bayesian estimation for linear Gaussian systems. Recursive: it repeats every instant and needs only the previous state, not the whole history. Bayesian: it combines a prior belief with evidence. Estimation: it infers something that cannot be seen. Linear: equations of the form Ax+bAx+bAx+b. Gaussian: uncertainty represented by normal distributions.

The mind does something qualitatively similar. Walking to meet someone last seen heading east at a steady pace, one predicts they should be at that corner by now; a shape moves in the fog, and the eyes are unreliable, so one does not conclude they are exactly where something moved. Prediction and a noisy glimpse are combined. When the fog lifts, sight counts more; when it thickens, the prediction does.

Outside the linear Gaussian case the rule still holds, but the posterior no longer fits in two numbers, and other filters carry it: a sum of Gaussians, a cloud of particles (Gaussian assumption).

Two voices, a model and a sensor, and a judge that keeps track of how much it knows about what it knows: that is state estimation in one line.