ML//evaluation//confusion matrix

A confusion matrix is the table that counts, for a classifier run at one threshold, how many cases fall in each combination of true class and predicted class, and every classification metric is a ratio of its cells. For a failure detector it has four: **true positives** (failures it flagged), **false positives** (false alarms), **false negatives** (failures it missed) and **true negatives** (healthy cases it left alone).


A confusion matrix is the table that counts, for a classifier run at one threshold, how many cases fall in each combination of true class and predicted class, and every classification metric is a ratio of its cells. For a failure detector it has four: true positives (failures it flagged), false positives (false alarms), false negatives (failures it missed) and true negatives (healthy cases it left alone).

A month of a fan monitor gives a concrete one. Of 10,000 one-hour windows, 10 preceded a real failure; the model raised 50 alarms, 8 of them before real failures. So TP=8TP=8TP=8, FP=42FP=42FP=42, FN=2FN=2FN=2 and TN=9948TN=9948TN=9948. Accuracy is (8+9948)/10000=99.6 %(8+9948)/10000=99.6,%(8+9948)/10000=99.6% and says almost nothing. Precision is 8/50=16 %8/50=16,%8/50=16% and recall 8/10=80 %8/10=80,%8/10=80%, and those two say what the crew will live with: about five visits for each failure caught, and one failure in five missed.

precision=TPTP+FP,recall=TPTP+FN,false alarm rate=FPFP+TN\text{precision}=\frac{TP}{TP+FP},\qquad \text{recall}=\frac{TP}{TP+FN},\qquad \text{false alarm rate}=\frac{FP}{FP+TN}precision=TP+FPTP​,recall=TP+FNTP​,false alarm rate=FP+TNFP​

These three ratios are behind precision and recall and, swept over every threshold, behind PR-AUC (precision against recall) and ROC-AUC (recall against false alarm rate).

It belongs to one threshold. Moving the threshold moves cases between columns, trading false alarms for misses, so the matrix worth showing is the one at the threshold the plant will actually run, and that threshold comes from costs (decision threshold).

Its cells are the statistician's errors under engineering names: a false positive is a type I error, a false negative a type II error (type I and type II errors). Multiplied cell by cell by the cost of each outcome, the matrix becomes the expected cost of running the detector, the number a decision actually needs.

With more than two classes it grows to one row and one column per class, and the off-diagonal cells show which classes the model confuses with which (a bearing defect read as misalignment), often more useful than any single score.