mathematics//decision theory//decision threshold
A decision threshold is the probability above which a system should act (raise an alarm, reject a part, stop a machine), and its right value follows from the costs of the two possible errors, which the classifier knows nothing about. It is the line between a model's output and an action in every predictive maintenance system, inspection camera and alarm, and the place where most of the money of such a system is won or lost.
A decision threshold is the probability above which a system should act (raise an alarm, reject a part, stop a machine), and its right value follows from the costs of the two possible errors, which the classifier knows nothing about. It is the line between a model's output and an action in every predictive maintenance system, inspection camera and alarm, and the place where most of the money of such a system is won or lost.
With ppp the probability of failure the model gives, CFPC_{FP}CFP the cost of a false alarm and CFNC_{FN}CFN the cost of a missed failure (correct decisions costing nothing), acting pays when staying silent costs more on average, p CFN>(1−p) CFPp,C_{FN} > (1-p),C_{FP}pCFN>(1−p)CFP, that is when
p>p∗=CFPCFP+CFN.p > p^{*} = \frac{C_{FP}}{C_{FP} + C_{FN}}.p>p∗=CFP+CFNCFP.
An inspection costs 500 euros and an undetected bearing failure 40,000, so p∗≈0.012p^*\approx0.012p∗≈0.012: it is worth inspecting at a 1.2% chance of failure. The cut sits far from 0.5 whenever the two errors cost differently, which in engineering is nearly always.
The threshold is a cost decision, and it only works with calibrated probabilities.
The classifier supplies ppp; the cost table supplies the cut, which moves when the costs change, with nothing retrained.
The 0.5 threshold trap: coming from machine learning, it is tempting to train a classifier, maximize accuracy and cut at 0.5. That cut is optimal only when both errors cost the same, and on an inspection line it loses money on every doubtful part.
The rule needs ppp to be a real frequency: of the cases the model scores at 10%, about 10% must fail. An overconfident network shifts every effective threshold, so the cost-optimal cut becomes a wrong one while accuracy still looks fine; reliability diagrams check it and isotonic or Platt scaling fix it (probability calibration). The rarer the failure, the more the base rate matters in what ppp can be.
It is distinct from the detection threshold of fault diagnosis, which is set on a residual statistic to fix a false-alarm rate when no cost table or failure probability is available; both trade false alarms against misses, and a ROC curve shows the menu either one picks from. This is the simplest case of decision theory, where the same weighing extends to more than two actions.