control//fault diagnosis//detection threshold
A detection threshold is the simplest fault detector, a fixed limit on a residual that raises an alarm when \(|r_k|>h\), and it is what most alarms in plants, cars and autopilots actually are: a high-temperature trip on a bearing, a low-pressure alarm on a pump, a gate on a GPS fix. Its appeal is that anyone can read it, it fits in a PLC, and when it fires the operator knows why.
A detection threshold is the simplest fault detector, a fixed limit on a residual that raises an alarm when ∣rk∣>h|r_k|>h∣rk∣>h, and it is what most alarms in plants, cars and autopilots actually are: a high-temperature trip on a bearing, a low-pressure alarm on a pump, a gate on a GPS fix. Its appeal is that anyone can read it, it fits in a PLC, and when it fires the operator knows why.
Its weakness is arithmetic. With a Gaussian residual of standard deviation σ\sigmaσ and h=3σh=3\sigmah=3σ, the probability of a false alarm on each sample is 0.27 %. That sounds small; sampled at 100 Hz it is almost a thousand false alarms an hour. False alarms are counted per hour, the false alarm rate per hour, because what runs out is the operator's patience: alarm management guidance aims at about one alarm every ten minutes per operator (alarm management), and an alarm that sounds all day ends up silenced, at which point it detects nothing.
The cheap remedies trade something for fewer alarms:
Raising hhh keeps only large faults visible.
Averaging NNN samples divides the noise by N\sqrt NN and adds a delay of about half the window.
Persistence logic requires the limit to be exceeded repeatedly (three of the last five samples, or continuously for ten seconds), which suppresses isolated spikes from a loose connector or a bit flip on a bus.
Hysteresis uses two limits, raising the alarm above honh_{\text{on}}hon and clearing it only below a lower hoffh_{\text{off}}hoff, so a residual hovering at the limit does not make the alarm chatter.
Beyond these, the elegant remedy is to accumulate evidence over time (CUSUM).
The threshold is a decision about costs.
It depends on how often faults occur and on what each error costs, so the same detector needs one limit on a standby pump and another on a critical compressor (decision threshold, and the cost curve of change detection).
The figure of 0.27 % for three sigma holds only for a residual that is Gaussian, zero-mean and white. With heavy tails the real rate is many times higher; with colored noise the residual runs long on one side and trips persistence rules and CUSUMs alike. Whiten the residual first, then calibrate the limit on real healthy data rather than on the formula.
A threshold with hysteresis is often all that remains of a machine-learning project once the right physical feature has been found: a limit on a pump's speed-normalized power, maintained by the technician on shift, beats a model nobody on site can retrain (physics-based features, flyswatter rule).