mathematics//statistics//hypothesis testing//type I and type II errors

Type I and type II errors are the two ways a decision rule on noisy data can fail: a **type I error** (false alarm, false positive) acts when nothing was happening, and a **type II error** (missed detection, false negative) fails to act when something was. Every detector, from a PLC threshold on a bearing temperature to a fault classifier, is tuned by choosing how many of each it will accept. Their rates are written \(\alpha\) (false alarms per test on healthy data) and \(\beta\) (misses per real event); \(1-\beta\) is the detector's **power**, or sensitivity.


Type I and type II errors are the two ways a decision rule on noisy data can fail: a type I error (false alarm, false positive) acts when nothing was happening, and a type II error (missed detection, false negative) fails to act when something was. Every detector, from a PLC threshold on a bearing temperature to a fault classifier, is tuned by choosing how many of each it will accept. Their rates are written α\alphaα (false alarms per test on healthy data) and β\betaβ (misses per real event); 1−β1-\beta1−β is the detector's power, or sensitivity.

Picture the statistic a detector watches as two overlapping bells, one for healthy readings and one for faulty ones, and the threshold as a vertical line between them. The healthy bell's tail beyond the line is α\alphaα; the faulty bell's tail on the near side is β\betaβ. Slide the line right and false alarms fall while misses grow; slide it left and the reverse. No threshold minimises both. Only separating the bells does that: a better sensor, a better model behind the residual, or more readings averaged per decision.

The rates per test compound with how often the test runs.

A 3σ3\sigma3σ threshold gives α=0.27%\alpha=0.27%α=0.27% per sample, which sounds small and is almost a thousand false alarms per hour at 100 Hz. A detector is specified in false alarms per hour or per year, and its α\alphaα per sample is derived from that.

The full trade-off of a detector is its ROC curve, the miss and false alarm rates of every possible threshold; one point on it is chosen by the costs (decision threshold). With a 2,000 € inspection against a 200,000 € breakdown, a false alarm is a hundred times cheaper than a miss, and the line belongs far to the left.

A third axis appears when evidence accumulates over time: detection delay. Methods that wait for several readings (CUSUM, persistence rules) buy fewer false alarms and fewer misses at the price of reacting later, which matters for a fast fault and costs nothing for a slow wear trend.

With rare faults the false alarm rate dominates what the operator sees. A detector with 99 % power and 1 % false alarms, on a fault present in one check in a thousand, produces eleven alarms per real fault (base rate fallacy); in that regime, lowering α\alphaα is worth far more than raising power.