mathematics//statistics//estimation

Estimation is the part of statistics that computes, from finite noisy data, a value for an unknown parameter of a model together with how wrong that value is likely to be, and it is what an engineer does when calibrating a sensor against a reference, fitting a motor's thrust curve or reading a pump's mean vibration from 25 measurements. The thing produced is an estimate, an opinion about a quantity nobody sees; the rule that produces it from data is an **estimator**, and estimation is the study of how good that rule is. When the unknown changes in time and must be followed tick by tick, the same ideas run recursively as state estimation.


Estimation is the part of statistics that computes, from finite noisy data, a value for an unknown parameter of a model together with how wrong that value is likely to be, and it is what an engineer does when calibrating a sensor against a reference, fitting a motor's thrust curve or reading a pump's mean vibration from 25 measurements. The thing produced is an estimate, an opinion about a quantity nobody sees; the rule that produces it from data is an estimator, and estimation is the study of how good that rule is. When the unknown changes in time and must be followed tick by tick, the same ideas run recursively as state estimation.

Every estimator errs in two ways of different nature, and they behave differently as data grow. The variance comes from having few data: repeat the test and the number moves, and it shrinks as 1/N1/\sqrt{N}1/N​ (standard error). The bias comes from a wrong model, data selected crookedly or a sensor out of calibration, and it does not shrink at all: a million readings of a thermocouple that sits 2 °C high give a very precise mean that is 2 °C wrong (precision and accuracy).

Least squares, maximum likelihood and MAP are one idea under three sets of assumptions.

Write how probable the data are for each candidate value (likelihood function); pick the value that makes them least surprising (maximum likelihood estimation), which under Gaussian, independent, equal-variance noise is exactly least squares; add what was known before measuring and the same fit becomes MAP estimation, which is regularization.

Maximum likelihood is the default when data are plentiful and the noise model is believable. With few points it believes the noise: four calibration points between 20 and 30 °C can return a gain of 0.83 for a sensor whose datasheet guarantees 1±0.051\pm0.051±0.05, and the prior of MAP is what pulls it back.

A number without its spread is half an answer. The standard error says how much the estimate would move with other data, and the confidence interval turns it into a range; for a quantity with no formula (a percentile, a ratio), the bootstrap gives the spread by resampling.

MLE and MAP both return one point. When the decision needs the width of the belief (how likely is the gain to be below 0.95?), the step up is Bayesian inference, which keeps the whole posterior.

The error travels downstream. A sensor σ\sigmaσ estimated wrong becomes a wrong measurement noise RRR in a Kalman filter, which then believes it knows more (or less) than it does, and from there a detection threshold set in the wrong place.