forecasting//forecast calibration

Forecast calibration is the agreement between the confidence attached to forecasts and the frequency with which they come true, and it is how a forecaster's probabilities are checked over many claims. A forecaster is calibrated if, among all the claims made at 80%, about eight in ten happen; a single forecast can never show it.


Forecast calibration is the agreement between the confidence attached to forecasts and the frequency with which they come true, and it is how a forecaster's probabilities are checked over many claims. A forecaster is calibrated if, among all the claims made at 80%, about eight in ten happen; a single forecast can never show it.

A common score is the Brier score, the mean squared gap between each probability pip_ipi​ and its outcome oio_ioi​ (1 if it happened, 0 if not):

BS=1N∑i=1N(pi−oi)2\mathrm{BS} = \frac{1}{N}\sum_{i=1}^{N} (p_i - o_i)^2BS=N1​i=1∑N​(pi​−oi​)2

Lower is better; always answering 50% scores 0.25.

The usual failure is overconfidence: intervals too narrow and probabilities too close to 0 or 1. One mechanical cause is treating sources that share data as independent, which adds their information twice (inverse-variance weighting assumes independence).

Calibration says nothing about how informative the forecasts are. Someone who always forecasts the base rate can be perfectly calibrated and useless; the second quality, sharpness, measures how far the forecasts dare to move from it.