ML//anomaly detection//isolation forest
An isolation forest is an anomaly detector that grows many random trees, each splitting the data on a random feature at a random value, and scores every point by how few splits it takes to isolate it; it is the usual first detector for unlabeled tables such as a plant's sensor snapshots, because it needs no distance, no scaling and little tuning. Liu, Ting and Zhou proposed it in 2008.
An isolation forest is an anomaly detector that grows many random trees, each splitting the data on a random feature at a random value, and scores every point by how few splits it takes to isolate it; it is the usual first detector for unlabeled tables such as a plant's sensor snapshots, because it needs no distance, no scaling and little tuning. Liu, Ting and Zhou proposed it in 2008.
The intuition is a game of twenty questions played with random cuts. A point in the middle of the dense cloud of normal operation needs many cuts before it stands alone; a point with an odd value, or an odd combination of values, falls into a box of its own after a few. Averaged over a few hundred trees, the path length becomes a stable score, and short paths mean rare. Each tree sees only a small subsample (256 points in the original paper), so training takes seconds and scoring a row costs about a thousand comparisons (a hundred trees, ten levels each), cheap enough for an industrial PC or an edge gateway.
Its input is a table of features, one row per time window or per unit; its output is a score per row, which still needs a threshold. That threshold comes from the number of alarms the team can review each week, or from a percentile on held-out healthy data. The library's contamination setting is a guess at the anomaly rate, and a wrong guess moves every alarm.
It finds global oddities well and contextual ones poorly. A temperature that is normal at full load and rare at no load looks ordinary to it unless load is a feature or the data are split by operating regime; residuals from a model of what is expected do better there (anomaly detection).
It is easily confused with a random forest, with which it shares only the trees. A random forest is supervised and chooses each split to separate labelled classes; an isolation forest has no labels and chooses its splits at random on purpose. Next to the Mahalanobis distance it makes no Gaussian assumption and copes with several clusters of normal behaviour; next to PCA monitoring it gives no contribution plot pointing at the broken correlation.
Like every detector it calls rare what was rare in its training window, so failures left in that window are learned as normal. Known incidents are removed from the training data first.