mathematics//statistics//Bayesian inference
Bayesian inference is the approach to statistics that treats unknown parameters as quantities with a probability distribution, updates that distribution with data by Bayes' rule, and reports the whole resulting posterior instead of one value, and engineers use it where data are scarce and prior knowledge is real: failure rates from a handful of failures, wear parameters of a fleet, a simulator calibrated against a few tests. Its output is a distribution over the parameter, from which a mean, a **credible interval** (the parameter lies here with 95 % probability) or the probability of any event that matters to a decision can be read.
Bayesian inference is the approach to statistics that treats unknown parameters as quantities with a probability distribution, updates that distribution with data by Bayes' rule, and reports the whole resulting posterior instead of one value, and engineers use it where data are scarce and prior knowledge is real: failure rates from a handful of failures, wear parameters of a fleet, a simulator calibrated against a few tests. Its output is a distribution over the parameter, from which a mean, a credible interval (the parameter lies here with 95 % probability) or the probability of any event that matters to a decision can be read.
The contrast with the rest of estimation is the width. Maximum likelihood returns the peak of the likelihood and MAP estimation the peak of the posterior: one number each, MAP being Bayes halfway. A maintenance decision usually turns on the tail (what is the probability that this bearing's life is below the next shutdown?), and only the full posterior answers it.
Writing the posterior is easy, p(θ∣y)∝p(y∣θ) p(θ)p(\theta\mid y)\propto p(y\mid\theta),p(\theta)p(θ∣y)∝p(y∣θ)p(θ); computing it is not, because the normalising constant p(y)p(y)p(y) is an integral over every possible θ\thetaθ. There are four routes, from cheapest to dearest, and choosing among them is most of the engineering.
A closed form exists when prior and likelihood belong to matching families (a conjugate prior). Gaussians with a linear model give a Gaussian, which is the Kalman filter. For a failure rate, a Beta(α,β)(\alpha,\beta)(α,β) prior and kkk failures in nnn trials give a Beta(α+k,β+n−k)(\alpha+k,\beta+n-k)(α+k,β+n−k) posterior: two additions in a spreadsheet, and it works with three failures.
With one to three parameters a grid approximation is enough: evaluate likelihood times prior on a grid of candidate values, divide by the sum, and read anything off the result. Its cost grows as the grid size to the power of the number of parameters, which is what kills it beyond three.
When the posterior is close to a bell, the Laplace approximation fits a Gaussian at the posterior peak, with a width from the curvature there; it costs one optimisation and gives almost the same answer as sampling.
Otherwise the posterior is represented by samples: online, a cloud of weighted hypotheses that moves with each reading (particle filter); offline, with a complicated model and time, MCMC, at thousands to millions of model evaluations.
Today's posterior is tomorrow's prior.
When new data arrive, the update starts from what is already believed, so inference can run one reading at a time, and estimating in real time is that, repeated hundreds of times a second.
The prior is the price and the benefit. With many data it washes out and the answer matches maximum likelihood; with few, it carries the physics or the fleet's history (random-effects pooling), and a wrong confident prior biases the answer until the data outvote it. Uncertainty quantification for deep networks (Bayesian neural networks, calibrated ensembles) is mostly research; closed forms and MCMC for reliability and calibration are standing practice.