ML//generalization//out-of-distribution

Out-of-distribution describes an input unlike those a model was trained on (new lighting on an inspection camera, a replaced sensor, fog, another machine, a heavier parcel), and the failure that comes with it: the model answers with its usual confidence, because nothing in it measures how familiar the input is. A network does not know that it does not know. (Two neighbouring terms mean other things: data drift is the production data slowly moving away from the training data, the process that brings such inputs, and distributional shift in language models is the change of next-token probabilities caused by the context.)


Out-of-distribution describes an input unlike those a model was trained on (new lighting on an inspection camera, a replaced sensor, fog, another machine, a heavier parcel), and the failure that comes with it: the model answers with its usual confidence, because nothing in it measures how familiar the input is. A network does not know that it does not know. (Two neighbouring terms mean other things: data drift is the production data slowly moving away from the training data, the process that brings such inputs, and distributional shift in language models is the change of next-token probabilities caused by the context.)

A learned model learns the physics of the data it saw, and only that. A drone model trained on calm days has never seen wind: inside its data it may beat any physical model, outside it can do anything, while a grey-box model degrades gently because its physics still holds (grey-box model). Confidence is no guide there, since a softmax output can be well calibrated on validation data and mean nothing on an input from elsewhere (probability calibration).

Nothing inside a network reports that it has left its distribution, so the distribution has to be watched from outside.

In a closed loop the stakes rise: a learned controller's small error carries the system into states it never saw, where it errs more (imitation learning). Critical systems therefore wrap the network in a monitor of its inputs, limits on its outputs and a verified classical fallback (safety filter).

An out-of-distribution detector checks whether each input lies within the training distribution: a Mahalanobis distance on the network's internal features, the reconstruction error of an autoencoder, or plain range checks on the raw inputs, which catch more in practice than their simplicity suggests.

Deep ensembles train several networks from different random seeds. Where they agree the input is familiar, and their disagreement is a useful uncertainty signal, paid for with several times the compute.

Uncertainty quantification in deep networks (Bayesian networks, calibrated ensembles) is mostly research with niche production uses. When real uncertainty is needed, a model that gives it by construction, a Gaussian process or a Kalman filter, is often the better engineering choice.