mathematics//statistics//estimation//maximum likelihood estimation
Maximum likelihood estimation is a method of fitting a model's parameters that picks the values under which the observed data would have been most probable, and it is the default way to calibrate sensors, fit failure distributions and train classifiers: choose the parameters that are least surprised by the data. It takes a model of how the data are generated, noise included, and the data; it returns a point, \(\hat\theta=\arg\max_\theta L(\theta)\), the peak of the likelihood function.
Maximum likelihood estimation is a method of fitting a model's parameters that picks the values under which the observed data would have been most probable, and it is the default way to calibrate sensors, fit failure distributions and train classifiers: choose the parameters that are least surprised by the data. It takes a model of how the data are generated, noise included, and the data; it returns a point, θ^=argmaxθL(θ)\hat\theta=\arg\max_\theta L(\theta)θ^=argmaxθL(θ), the peak of the likelihood function.
The real case shows why it matters. A cheap temperature sensor is calibrated against a reference with the model yi=aTi+b+eiy_i=aT_i+b+e_iyi=aTi+b+ei, where eie_iei is Gaussian, independent, zero-mean and of the same variance σ2\sigma^2σ2 at every point. Taking logs,
−logL(a,b)=12σ2∑i=1N(yi−aTi−b)2+Nlog (2π σ).-\log L(a,b)=\frac{1}{2\sigma^2}\sum_{i=1}^{N}\left(y_i-aT_i-b\right)^2+N\log\!\left(\sqrt{2\pi}\,\sigma\right).−logL(a,b)=2σ21i=1∑N(yi−aTi−b)2+Nlog(2πσ).
yiy_iyi is the reading, TiT_iTi the reference temperature, aaa the gain and bbb the offset. The second term does not involve aaa or bbb, so maximising the likelihood is minimising the sum of squares: least squares is maximum likelihood with Gaussian, independent, constant-variance noise.
Choosing a loss is choosing a noise model, whether or not one meant to.
Give each point its own variance σi2\sigma_i^2σi2 and the fit becomes weighted least squares with weights 1/σi21/\sigma_i^21/σi2, the same weights as inverse-variance weighting in sensor fusion. Assume noise with more frequent spikes than a Gaussian (a Laplace distribution) and it becomes the sum of absolute errors, whose minimum in the simplest case is the median (robust statistics). Assume a yes-or-no outcome and it becomes the cross-entropy of logistic regression.
With enough data it is hard to beat. Under mild conditions the estimate converges to the true value and its spread approaches the smallest any unbiased method can reach, and the curvature of the log-likelihood at the peak gives its standard error for free.
With few data it believes the noise. Four calibration points squeezed between 20 and 30 °C can return a gain of 0.83 when the datasheet guarantees 1±0.051\pm0.051±0.05; adding what was known before measuring is MAP estimation.
It is only as right as the noise model it encodes. One absurd reading from a flipped bit drags a Gaussian fit as far as it likes, and lifetimes cut short by preventive replacement bias a fit that treats them as failures (censored data).
Closed forms are the exception. Logistic regression, a Weibull fit or a filter's noise levels are maximised numerically, usually with gradient descent or Newton steps on the negative log-likelihood.