mathematics//probability//expected value
The expected value of a random variable is the average of its possible values weighted by their probabilities, and it is the single number used to summarize where a quantity sits and to rank decisions by their average cost. It is the centre of mass of the probability distribution: put the probability of each value on a rod as a weight and the expected value is the point where the rod balances.
The expected value of a random variable is the average of its possible values weighted by their probabilities, and it is the single number used to summarize where a quantity sits and to rank decisions by their average cost. It is the centre of mass of the probability distribution: put the probability of each value on a rod as a weight and the expected value is the point where the rod balances.
E[X]=∑xx p(x)orE[X]=∫x p(x) dx\mathbb E[X]=\sum_x x\,p(x)\qquad\text{or}\qquad \mathbb E[X]=\int x\,p(x)\,dxE[X]=x∑xp(x)orE[X]=∫xp(x)dx
The sum is for discrete values, the integral for continuous ones, and both are written μ\muμ, the mean. A gyroscope lying still on a bench reads on average 0.30.30.3°/s with a spread of 0.050.050.05°/s around it. The true rate is zero, so the nonzero mean of its error is a sensor bias, a systematic error that calibration measures and subtracts; the spread, the variance around the mean, is noise, which averaging shrinks and nothing removes.
Expectation passes through sums and stops at curves.
Linearity, E[aX+bY]=a E[X]+b E[Y]\mathbb E[aX+bY]=a,\mathbb E[X]+b,\mathbb E[Y]E[aX+bY]=aE[X]+bE[Y] holds always, independent or not, which is why linear filters can propagate means exactly. For a nonlinear function E[f(X)]≠f(E[X])\mathbb E[f(X)]\neq f(\mathbb E[X])E[f(X)]=f(E[X]): the power in the wind goes as the cube of its speed, so a turbine sized on the mean wind speed underestimates the energy, and a filter that pushes only the mean through a curve is biased.
The expected value of a cost is what decision theory minimizes: each action is scored by its average cost over the states of the world, weighted by how credible each state is. It treats many small losses and one total loss alike, which is right for scrap and wrong for a drone falling on people.
From data, the expected value is estimated by the sample mean, whose error shrinks as σ/N\sigma/\sqrt Nσ/N (standard error). That works when the variance is finite. With heavy-tailed distributions the sample mean converges slowly, and the Cauchy distribution has no mean at all: averaging a million of its samples is no better than reading one.
The mean is one summary among several. In a skewed distribution (repair times, latencies) it sits well above the typical value and below the bad cases, which is why systems that must meet deadlines are judged by a percentile.