industrial//industrial automation//alarm management

Alarm management is the engineering discipline of designing, rationalizing and maintaining the alarms a control system raises to its operators, so that each alarm is real, timely and calls for an action; it exists because a SCADA screen with thousands of possible alarms is only useful if the few that sound at once can be acted on. The reference guidance, ISA-18.2 and EEMUA 191, puts a manageable rate at around one alarm every ten minutes per operator in normal operation.


Alarm management is the engineering discipline of designing, rationalizing and maintaining the alarms a control system raises to its operators, so that each alarm is real, timely and calls for an action; it exists because a SCADA screen with thousands of possible alarms is only useful if the few that sound at once can be acted on. The reference guidance, ISA-18.2 and EEMUA 191, puts a manageable rate at around one alarm every ten minutes per operator in normal operation.

That target collides with statistics. A limit at 3σ3\sigma3σ on a Gaussian signal raises a false alarm on 0.27 % of checks; across a thousand sensors checked once a second that is 2.7 spurious alarms at every look and more than two hundred thousand a day, the multiple-comparisons trap of hypothesis testing. The result has a name, alarm fatigue: operators learn that the horn means nothing, acknowledge without reading, and eventually silence it, and a silenced detector detects zero faults. The classic case is the 1994 explosion at the Texaco refinery in Milford Haven, where the two operators had to recognize and act on 275 alarms in the last eleven minutes, most of them displayed as high priority1.

1UK Health and Safety Executive, The explosion and fires at the Texaco Refinery, Milford Haven, 24th July 1994.

A robust threshold monitor for one variable.

Take a clean week of data, its median, a robust spread (the median absolute deviation, scaled to σ\sigmaσ), a limit at 4σ4\sigma4σ and a persistence rule. It fits in a PLC, it solves most single-variable cases, and any operator understands why it fired; cost-based thresholds earn their extra complexity only when the costs are very asymmetric and the base rate is known.

The cheap rules that cut a flood are those of the detection threshold: persistence (five samples in a row outside), hysteresis (separate raise and clear limits so an alarm does not chatter), and detectors that accumulate evidence such as CUSUM, which replace many twitchy limits with one patient one.

Rationalization is the organizational half: every alarm gets a documented cause, consequence, operator action and priority, and an alarm with no action is removed or demoted to an event log. Alarms that are always on (standing alarms) and alarms that toggle within seconds (chattering) are the first targets.

Most alarms are false even for a good detector when real faults are rare, because the false alarms are drawn from the much larger healthy population (base rate fallacy); the number to manage is alarms per hour, and the number to report is the share of alarms that led to an action.