ML//evaluation//precision and recall

Precision and recall are the two classification metrics that read a detector in the language of alarms: precision is the share of the alarms raised that were real, recall (also called sensitivity, or detection rate) the share of the real events that raised an alarm, and together they are the honest pair for any problem where the interesting class is rare, such as failures, defects or intrusions. Precision answers the maintenance crew (*how many of these call-outs are wasted?*); recall answers the plant manager (*how many failures will still take us by surprise?*). (The *precision* of a measuring instrument, its repeatability, is a different notion: precision and accuracy.)


Precision and recall are the two classification metrics that read a detector in the language of alarms: precision is the share of the alarms raised that were real, recall (also called sensitivity, or detection rate) the share of the real events that raised an alarm, and together they are the honest pair for any problem where the interesting class is rare, such as failures, defects or intrusions. Precision answers the maintenance crew (how many of these call-outs are wasted?); recall answers the plant manager (how many failures will still take us by surprise?). (The precision of a measuring instrument, its repeatability, is a different notion: precision and accuracy.)

They pull against each other through the threshold. Lowering it raises recall, since more true events cross it, and lowers precision, since more healthy cases cross it too; the curve traced by sweeping the threshold is summarized by PR-AUC. Neither looks at the true negatives, which is exactly why they survive class imbalance: thousands of healthy windows correctly ignored cannot inflate them, as they inflate accuracy and the false alarm rate behind ROC-AUC.

A low precision can be the right answer. When a miss costs forty times an inspection, a detector whose alarms are real one time in ten pays for itself. The costs fix the threshold, and precision and recall are then read at that threshold, never maximized for their own sake (decision threshold).

The F1 score, their harmonic mean 2PR/(P+R)2PR/(P+R)2PR/(P+R), is pulled toward the smaller of the two (it is easy to misremember as their sum or their plain average: a detector with perfect precision and 10 % recall scores 0.18, while the average would flatter it at 0.55). It weighs them equally, which almost never matches a plant where a miss and a false alarm cost very different amounts. It is a convenience for ranking models and a poor basis for choosing a threshold.

Precision depends on the base rate, and recall does not. The same detector moved from a fleet where 1 % of units fail to one where 0.1 % fail keeps its recall and loses most of its precision, because the false alarms stay while the true events thin out (base rate). Both come from the counts of the confusion matrix.