mathematics//statistics//hypothesis testing

Hypothesis testing is a statistical procedure that computes a number from data and compares it with a threshold to decide whether to reject a default assumption (the **null hypothesis**, usually *nothing is happening*), and in engineering it is the machinery under every alarm, acceptance test and A/B comparison: a decision rule in disguise. Compute a statistic (a residual in sigmas, a vibration level, a difference of means), and act if it crosses the line.


Hypothesis testing is a statistical procedure that computes a number from data and compares it with a threshold to decide whether to reject a default assumption (the null hypothesis, usually nothing is happening), and in engineering it is the machinery under every alarm, acceptance test and A/B comparison: a decision rule in disguise. Compute a statistic (a residual in sigmas, a vibration level, a difference of means), and act if it crosses the line.

The rule can be wrong in two ways, acting when nothing was happening and missing what was, and moving the threshold trades one for the other (type I and type II errors). The significance level is the false alarm rate accepted per test, by habit 5 %. That habit has no engineering content. A threshold is a decision and belongs where the costs say: with a 2,000 € inspection and a 200,000 € compressor failure, inspect whenever the probability of fault exceeds about 1 % (decision threshold).

The p-value is the probability of data this extreme if nothing were happening.

It is a different number from the probability that nothing is happening, and reading one as the other is the base rate fallacy again: a p-value of 0.01 on a machine that almost never fails can still leave a fault unlikely.

Many tests multiply the false alarms. A thousand sensors each with a threshold at 3σ3\sigma3σ (0.27 % false alarms per check, if the noise were Gaussian and independent) give 2.7 spurious alarms every time they are read; checked once a second, more than 200,000 a day. Industrial alarm-management guidance (ISA-18.2, EEMUA 191) aims at the order of one alarm per ten minutes per operator in normal operation (alarm management).

The gap is closed by asking for more evidence before acting. Persistence (five samples in a row outside), hysteresis (a lower line to clear the alarm than to raise it) and detectors that accumulate evidence over time (CUSUM) each cut false alarms by orders of magnitude for a small delay.

The test inherits every assumption of its statistic. Thresholds in sigmas assume Gaussian, independent noise; real noise has heavier tails and memory, so a nominal 0.27 % becomes several percent (heavy-tailed distribution, effective sample size).

Flyswatter first26 · oct 09Read noteTo watch one variable, a clean week of data, a mean, a robust standard deviation, a 4σ4\sigma4σ threshold and a persistence rule solve most cases. It fits in a PLC and any operator understands why it fired; explicit cost-based decision theory pays only when costs are very asymmetric, the base rate is known and an alarm triggers something expensive (flyswatter rule).