mathematics//statistics//estimation//MAP estimation
MAP estimation (maximum a posteriori) is a method of fitting parameters that maximises the likelihood of the data multiplied by a prior belief about the parameters, and it is used whenever something is known before measuring (a datasheet value, a fleet's typical gain) and the data alone are too few to trust. It returns the peak of the posterior of Bayes' rule, \(p(\text{data}\mid\theta)\,p(\theta)\), where maximum likelihood estimation returns the peak of the first factor only.
MAP estimation (maximum a posteriori) is a method of fitting parameters that maximises the likelihood of the data multiplied by a prior belief about the parameters, and it is used whenever something is known before measuring (a datasheet value, a fleet's typical gain) and the data alone are too few to trust. It returns the peak of the posterior of Bayes' rule, p(data∣θ) p(θ)p(\text{data}\mid\theta),p(\theta)p(data∣θ)p(θ), where maximum likelihood estimation returns the peak of the first factor only.
Take the sensor calibrated with four points between 20 and 30 °C, where maximum likelihood returned a gain of 0.83 and an offset of 4.5 °C. The datasheet says gain 1±0.051\pm0.051±0.05 and offset 0±20\pm20±2 °C. With a Gaussian prior centred on the nominal value θ0\theta_0θ0 with spread τ\tauτ, maximising the posterior is
θ^MAP=argminθ 1σ2∑i(yi−f(xi;θ))2+1τ2∥θ−θ0∥2.\hat\theta_{\text{MAP}}=\arg\min_\theta\;\frac{1}{\sigma^2}\sum_i\left(y_i-f(x_i;\theta)\right)^2+\frac{1}{\tau^2}\left\|\theta-\theta_0\right\|^2 .θ^MAP=argθminσ21i∑(yi−f(xi;θ))2+τ21∥θ−θ0∥2.
f(xi;θ)f(x_i;\theta)f(xi;θ) is the model (here aTi+baT_i+baTi+b), θ0\theta_0θ0 the prior opinion and τ\tauτ how far it is trusted. The answer becomes a gain of 0.98 and an offset of 1.2 °C, close to the truth of 0.97 and 1.2. With many data the first term dominates and the prior fades; with few, the prior keeps the fit from believing the noise.
Regularizing is having a prior and admitting it.
The penalty above is exactly the ridge, or Tikhonov, regularization of ridge regression, with λ=σ2/τ2\lambda=\sigma^2/\tau^2λ=σ2/τ2: a tight prior is a strong penalty. A Laplace prior gives the absolute-value penalty of lasso instead.
The prior can be implemented as invented measurements. Append to the least-squares system one row per parameter, with weight σ/τ\sigma/\tauσ/τ, saying that the parameter equals its nominal value; an ordinary least-squares solver then returns the MAP estimate, and the prior behaves like a few extra readings of the datasheet.
It is Bayes halfway. MAP still returns one number and throws away the width of the posterior, which a decision often needs; keeping it is Bayesian inference.
The prior must come from knowledge, never from the same data. A prior tuned until the fit looks right is a second free parameter, and a confident prior that is wrong (a datasheet for another model) biases the answer for as long as data are scarce.
In a running filter the same idea is the prediction acting as prior for each reading: a Kalman filter update is a MAP estimate with Gaussian prior and Gaussian likelihood.