mathematics//statistics//estimation//likelihood function
The likelihood function is the probability of the data actually observed, read as a function of the unknown parameters of a model, and it is the score that every fitting method in statistics maximises or weighs: it says how surprised each candidate model would be by the readings in hand. For parameters \(\theta\),
The likelihood function is the probability of the data actually observed, read as a function of the unknown parameters of a model, and it is the score that every fitting method in statistics maximises or weighs: it says how surprised each candidate model would be by the readings in hand. For parameters θ\thetaθ,
L(θ)=p(data∣θ).L(\theta)=p(\text{data}\mid\theta).L(θ)=p(data∣θ).
The data are fixed and θ\thetaθ moves. That reading is what makes it confusing: L(θ)L(\theta)L(θ) is no probability over θ\thetaθ (it need not integrate to one over the parameters), and a value of 0.3 means nothing by itself. Only ratios matter, so a likelihood compares candidates: if one gain makes the calibration data ten times more probable than another, the data favour it tenfold.
A small case makes it concrete. A batch of 20 valves is pressure-tested and 3 fail. For a failure probability ppp, the probability of that outcome is L(p)∝p3(1−p)17L(p)\propto p^3(1-p)^{17}L(p)∝p3(1−p)17; it is tiny for p=0.5p=0.5p=0.5, largest at p=3/20=0.15p=3/20=0.15p=3/20=0.15, and tiny again near zero. The peak is the maximum likelihood estimate, and the width of the hump around it is how much the 20 tests pin ppp down.
In practice one works with the log. Independent readings multiply their probabilities, which underflow a computer after a few hundred points, so the log-likelihood ℓ(θ)=∑ilogp(yi∣θ)\ell(\theta)=\sum_i\log p(y_i\mid\theta)ℓ(θ)=∑ilogp(yi∣θ) turns the product into a sum with the same maximum. Optimisers minimise its negative, which is why a loss function and a negative log-likelihood are so often the same object (cross-entropy is one).
The noise model writes the likelihood. With Gaussian noise of variance σ2\sigma^2σ2, −ℓ-\ell−ℓ is the sum of squared residuals over 2σ22\sigma^22σ2 plus a constant, and maximising it is least squares; with spikier noise the same step gives another loss (robust statistics).
It carries partial information without throwing it away. A bearing still running at 1,200 h has not failed, and its term in the likelihood is the probability of surviving that long instead of a density at a failure time; that is how censored data enter a fit.
Multiplied by a prior it becomes the numerator of Bayes' rule, so MAP estimation and Bayesian inference start from it; divided between two hypotheses it is the likelihood ratio a detector compares with a threshold (hypothesis testing).