ML//anomaly detection
Anomaly detection is the family of methods that learn what normal operation looks like and flag what departs from it, and it is used where failures are too rare to learn from directly: three bearing failures in five years of a plant's history, one fraud in ten thousand payments, an intrusion in a month of network traffic. It turns the question *what does a failure look like?* into *what does normal look like?*, which only needs plenty of healthy data.
Anomaly detection is the family of methods that learn what normal operation looks like and flag what departs from it, and it is used where failures are too rare to learn from directly: three bearing failures in five years of a plant's history, one fraud in ten thousand payments, an intrusion in a month of network traffic. It turns the question what does a failure look like? into what does normal look like?, which only needs plenty of healthy data.
Every detector, from the cheapest to the deepest, produces an anomaly score s(x)s(x)s(x) and compares it with a threshold; the score is a residual whose expected value was learned from past data instead of derived from physics. Anomalies come in three kinds. A point anomaly is a single odd value (a spike); a contextual anomaly is odd only given its context (80 °C is normal for a motor at full load and very rare at no load); a collective anomaly is a sequence that is abnormal although each of its points is normal.
A detector finds what is rare, and rare is not the same as bad.
A new product recipe, the first hot day of summer or a sensor swapped by maintenance are legitimate anomalies. An anomaly is rare relative to history; a fault is wrong relative to how the system should behave. The strongest setup combines the two: a model of what is expected, and an anomaly detector on its residual.
The methods form a ladder from cheap to expensive. Physical limit checks per variable come first (a thermocouple reading −40 °C inside an oven needs no ML). Then the Mahalanobis distance to normal, which catches odd combinations of individually normal values, and its plant form, multivariate statistical process control. Then isolation forest, the usual default on tables. Then the one-class SVM, a support vector machine boundary drawn around normal data, and the local outlier factor, which compares each point's density with its neighbours', both finer and costlier. The autoencoder closes the ladder for curved relations and raw signals.
Residual-based anomaly detection is the best tool for contextual anomalies: regress the expected value from its context on healthy data (motor temperature from load and ambient), watch the residual, and set limits per operating regime.
Without labels, a detector is evaluated with injected faults, historical incidents and a weekly review by maintenance of the twenty highest alarms; how many alarms a crew can bear is the business of alarm management.
The deep end pays off only with hundreds of different assets, failure modes with no physical model, abundant data and expensive failures. A clogging air filter is caught by the pressure drop corrected for flow (it grows roughly with the square of the flow in turbulent conditions) and a threshold, which fits in a PLC and is understood when it fires (flyswatter rule). After detection comes fault diagnosis: what is wrong, and where.