mathematics//statistics//sample size

Sample size is the number of independent observations behind an estimate, and working it out before collecting data is how an engineer knows whether a test campaign, a validation or a historian can answer the question at all. A handful of back-of-the-envelope rules saves whole projects, because the answer is often far from intuition in both directions: a mean needs a few hundred points, a rare failure rate needs millions of hours.


Sample size is the number of independent observations behind an estimate, and working it out before collecting data is how an engineer knows whether a test campaign, a validation or a historian can answer the question at all. A handful of back-of-the-envelope rules saves whole projects, because the answer is often far from intuition in both directions: a mean needs a few hundred points, a rare failure rate needs millions of hours.

For a mean, the error falls as σ/N\sigma/\sqrt Nσ/N​ (standard error). To know it within a margin EEE at 95 %, N=(1.96 σ/E)2N=(1.96,\sigma/E)^2N=(1.96σ/E)2; to 10 % of σ\sigmaσ, (1.96/0.1)2≈400(1.96/0.1)^2\approx400(1.96/0.1)2≈400 independent samples. Halving the margin quadruples the bill.

For rare events, the events count, never the samples. With kkk events observed, the relative error of the estimated rate is about 1/k1/\sqrt k1/k​: a hundred events for ±10%\pm10%±10%. Validating by test a detector that should give one false alarm every 10,000 hours, to ±10%\pm10%±10%, needs about a million hours of operation.

With zero events, the rule of three.

If nothing has happened in NNN independent trials, all that can be claimed at 95 % is that the rate is below 3/N3/N3/N. Three hundred flights without incident show a rate under 1 % per flight; they say nothing about it being zero. That is why rare failures are certified with models, accelerated tests and simulation as well as with kilometres (functional safety).

Correlated samples are worth less. A signal with memory sampled fast carries far fewer independent observations than rows: an hour at 1 kHz is 3.6 million samples, and with a second of memory about 1,800 independent ones (effective sample size).

Coverage comes before volume. A thousand hours in summer say nothing about winter, and a thousand failures without a reliable label say little about their causes (industrial labels, selection bias). The question is which operating conditions the sample spans, and the count only matters within each one.

Rare classes set the size of a learning dataset. With one fault in 2,000 windows, a dataset of 10,000 windows holds five faults, too few to train and almost nothing to validate; a model that always says normal scores 99.95 % accuracy on it (base rate).