mathematics//statistics//estimation//standard error

The standard error is the standard deviation of an estimate across the datasets that could have been collected, and it is the number that says how much a computed mean, rate or fitted gain would move if the measurement were repeated: the uncertainty of the estimate, as opposed to the spread of the readings. For the mean of \(N\) independent readings with spread \(\sigma\),


The standard error is the standard deviation of an estimate across the datasets that could have been collected, and it is the number that says how much a computed mean, rate or fitted gain would move if the measurement were repeated: the uncertainty of the estimate, as opposed to the spread of the readings. For the mean of NNN independent readings with spread σ\sigmaσ,

SE=σN.\mathrm{SE}=\frac{\sigma}{\sqrt{N}} .SE=N​σ​.

A pump's vibration velocity is measured 25 times: mean 4.2 mm/s, standard deviation 0.8 mm/s. The readings scatter by 0.8, but the mean is known to 0.8/5=0.160.8/5=0.160.8/5=0.16 mm/s. The two numbers answer different questions. σ\sigmaσ is how much the next reading will differ from the last; the standard error is how well the average of these 25 pins down the pump's true level, and it is what a confidence interval is built from.

The square root is a law of diminishing returns.

Halving the uncertainty of a mean costs four times the data, dividing it by ten costs a hundred times. Averaging is cheap at first and ruinous later, which is why a better sensor or a longer integration of the right signal often beats more samples.

It shrinks only the random part. A million readings of a thermocouple offset by 2 °C give a standard error of a few thousandths of a degree around a value that is 2 °C wrong: the formula measures variance, and bias is the business of sensor calibration (precision and accuracy).

It assumes independent readings. With memory between samples the NNN in the formula must be the effective sample size, often thousands of times smaller than the row count; a slowly drifting signal logged at 1 kHz has a far larger standard error than σ/N\sigma/\sqrt Nσ/N​ suggests.

Every estimator has one, with its own formula. For an estimated rate p^\hat pp^​ from NNN trials it is p^(1−p^)/N\sqrt{\hat p(1-\hat p)/N}p^​(1−p^​)/N​: with a fault rate of 0.1 % and 10,000 records there are about 10 faults, and the rate is known only to about a third of itself (sample size). For a fitted parameter it comes from the curvature of the log-likelihood (maximum likelihood estimation); for anything without a formula, from the bootstrap.

The mean of many readings is close to Gaussian even when the readings are not (central limit theorem), which is why intervals of a mean as estimate plus or minus two standard errors work so widely.