mathematics//probability//Bayes' rule//base rate fallacy

The base rate fallacy is the error of judging the probability of a cause from the evidence alone while ignoring how rare the cause is, confusing \(P(\text{alarm}\mid\text{fault})\) with \(P(\text{fault}\mid\text{alarm})\), and it is why most alarms in a plant are false even when the detector is good. Any detector is described by two rates measured on the bench: its **sensitivity** (the share of real faults it flags) and its **false alarm rate** (the share of healthy cases it flags anyway). What the operator needs is a third number neither of them gives: when it rings, how likely is it real.


The base rate fallacy is the error of judging the probability of a cause from the evidence alone while ignoring how rare the cause is, confusing P(alarm∣fault)P(\text{alarm}\mid\text{fault})P(alarm∣fault) with P(fault∣alarm)P(\text{fault}\mid\text{alarm})P(fault∣alarm), and it is why most alarms in a plant are false even when the detector is good. Any detector is described by two rates measured on the bench: its sensitivity (the share of real faults it flags) and its false alarm rate (the share of healthy cases it flags anyway). What the operator needs is a third number neither of them gives: when it rings, how likely is it real.

A fleet of compressors has a bearing-damage detector run once a day on each machine. From history, one evaluation in 1,000 corresponds to a damaged bearing. The detector catches 99 % of real faults and false-alarms on 1 % of healthy machines. It rings. Most people answer 99 %. Count cases instead: of 100,000 evaluations, 100 are real faults and the detector flags 99; the other 99,900 are healthy and 1 % of them ring, 999 false alarms. Of 1,098 alarms, 99 are real: 9 %, eleven alarms for every real fault from a detector that looked excellent in the lab. Done in general, that count is Bayes' rule with the base rate as prior.

P(fault | alarm)9.0 % alarms per real fault11 faults missed1.0 % Fault rate 0.10 %, sensitivity 99 %, false alarms 1.0 %. Out of every 1,000 hours: 1 with a real fault and 11 alarms. Of those alarms, 1 is a real fault and 10 are false. An alarm means a real fault 9.0 % of the time.

Read the share of alarms that are real, then push sensitivity to the maximum and see how little it changes; cut the false alarm rate tenfold and see how much; require a second independent detector and count the alarms per real fault.

With a low base rate, the false alarm rate matters far more than the sensitivity.

Raising sensitivity from 99 % to 99.9 % leaves the 9 % at 9 %; cutting false alarms from 1 % to 0.1 % lifts it to 50 %. Raise the fault rate to 20 % and the same detector is right 96 % of the time.

A second detector helps only if it is independent. An independent, equally good one multiplies the odds by 99 again and lifts the 9 % to 91 %; two algorithms reading the same accelerometer fail together when it comes loose (common-mode failure).

A detector worth keeping can still be wrong nine times in ten. With a 2,000 € inspection against a 200,000 € breakdown, a 9 % chance of fault is an excellent reason to inspect (decision threshold); the failure is to deploy it without having done the count.

The same confusion hides in statistics: a p-value is the probability of the data if nothing were happening, and reading it as the probability that nothing is happening is this fallacy again (hypothesis testing). A classifier's accuracy on a rare class misleads the same way, which is why rare events are judged by precision and recall (ROC-AUC).