control//fault diagnosis//change detection

Change detection is the online statistical problem of deciding, as each sample arrives, whether the statistics of a signal have changed (its mean jumped, a slow drift began, its spread grew), and it is the decision step behind fault alarms, statistical process control charts and the integrity monitors of navigation systems. Its input is usually a residual whose healthy behaviour is known; its output is an alarm and, ideally, an estimate of when the change happened.


Change detection is the online statistical problem of deciding, as each sample arrives, whether the statistics of a signal have changed (its mean jumped, a slow drift began, its spread grew), and it is the decision step behind fault alarms, statistical process control charts and the integrity monitors of navigation systems. Its input is usually a residual whose healthy behaviour is known; its output is an alarm and, ideally, an estimate of when the change happened.

It can be wrong in two ways, and they pull against each other. A false alarm stops a healthy line; a late alarm lets a real fault run. In a one-shot test the trade is drawn as a ROC curve, probability of detection against probability of false alarm as the threshold moves (type I and type II errors). A detector that runs forever needs a different currency, time: the mean time between false alarms on a healthy signal against the mean delay between a change and its alarm.

Fewer false alarms are cheap in delay.

For the CUSUM, the mean time between false alarms grows roughly exponentially with its threshold while the detection delay grows only linearly, so cutting false alarms tenfold costs a few samples of delay. The point on that curve is picked by costs, never by habit.

With cFAc_{FA}cFA​ the cost of each false alarm, λFA(h)\lambda_{FA}(h)λFA​(h) the false alarms per hour at threshold hhh, cdc_dcd​ the cost of each hour a real fault goes unnoticed, λf\lambda_fλf​ the real faults per hour and τˉ(h)\bar\tau(h)τˉ(h) the mean delay in hours, the expected cost per hour is approximately

J(h)=cFA λFA(h)+cd λf τˉ(h).J(h) = c_{FA}\,\lambda_{FA}(h) + c_d\,\lambda_f\,\bar\tau(h).J(h)=cFA​λFA​(h)+cd​λf​τˉ(h).

If each false alarm costs 500 euros (fifteen minutes of stopped line and a technician's visit) and faults are rare, the minimum sits at a high threshold; if a missed fault is dangerous or expensive, the threshold comes down and the technician visits more often. Every detector is one more case of error relocation: the error is moved between the two columns, never removed.

The instantaneous detection threshold (a Shewhart chart in quality control) is the simplest detector and the right one for large, sudden changes. Accumulating detectors (CUSUM, the EWMA chart) see small persistent shifts that a threshold never crosses; when the size of the change is unknown, the generalized likelihood ratio test estimates it on the fly at more cost per sample.

All of them assume the healthy residual is known and its samples independent. A residual with autocorrelation runs long on one side and is read as a change, and a residual that drifts with the operating point (hotter every summer) is a change in the weather; whiten and normalize before detecting.