ML//evaluation//probability calibration

Probability calibration is the property of a classifier whose predicted probabilities match observed frequencies (among all the cases it gives 10 %, about 10 % fail), and the set of methods that restore it after training; it is the prerequisite for setting a threshold from costs, because a cost threshold compares a probability with a ratio of euros, and a probability that does not mean what it says moves the cut. The same property, applied to forecasts of events, is forecast calibration.


Probability calibration is the property of a classifier whose predicted probabilities match observed frequencies (among all the cases it gives 10 %, about 10 % fail), and the set of methods that restore it after training; it is the prerequisite for setting a threshold from costs, because a cost threshold compares a probability with a ratio of euros, and a probability that does not mean what it says moves the cut. The same property, applied to forecasts of events, is forecast calibration.

It is checked with a reliability diagram: cases are grouped by predicted probability (0 to 10 %, 10 to 20 %, and so on), and the observed share of positives in each bin is plotted against the mean prediction. A calibrated model lies on the diagonal; a curve below it is overconfident, one above it underconfident. With a threshold at 1.2 % (an inspection costing 500 euros against a missed failure costing 40,000), an error of a few points in the lowest bins changes how many machines get inspected each week (decision threshold).

Overconfidence is the usual failure of large modern networks: when they say 99 %, they are right less often. A softmax output is positive and sums to one, so it looks like a probability, and it becomes one only once calibrated and checked (softmax). Training tricks against class imbalance (reweighting, resampling) distort probabilities the other way.

A language model's stated confidence is a separate case. When it writes I am 95 % sure, the number is generated text and carries no calibration by default, and even its internal token probabilities tend to lose calibration after RLHF. Before a stated confidence decides anything (accepting an answer, skipping a human check), it is measured on labelled cases or replaced by an external check, such as agreement between several independent samples or a test the answer must pass.

Calibration is repaired after training, on held-out validation data. Platt scaling fits a sigmoid to the model's scores; isotonic regression fits a monotone step function, more flexible and hungrier for data; temperature scaling divides a network's logits by one fitted constant, which partly recalibrates the softmax without changing which class wins.

It holds only for data like those it was measured on: outside the training distribution a calibrated probability means nothing (out-of-distribution). It also says nothing about sharpness, since a model that predicts the base rate for every case is perfectly calibrated and useless.

A miscalibrated probability is one of the homes of model error and the cheapest to repair; left alone, it becomes a misplaced threshold that stops healthy machines or lets failing ones run (applied ML).