ML//MLOps//data drift

Data drift is the gradual or sudden departure of the data a deployed model receives from the data it was trained on, and watching for it is how an operator learns that a model has started working outside what it knows, usually long before its accuracy can be measured. The world changes after training: a sensor is replaced, a new bearing supplier arrives, winter lowers every temperature in the plant, a machine wears. The model does not notice; its outputs simply become less trustworthy.


Data drift is the gradual or sudden departure of the data a deployed model receives from the data it was trained on, and watching for it is how an operator learns that a model has started working outside what it knows, usually long before its accuracy can be measured. The world changes after training: a sensor is replaced, a new bearing supplier arrives, winter lowers every temperature in the plant, a machine wears. The model does not notice; its outputs simply become less trustworthy.

Two kinds are worth telling apart, because they call for different responses. In covariate shift the inputs move while the rule linking them to the output stays the same (a new sensor supplier with a different offset, cold mornings): the model is being asked about a region it saw little of. In concept drift the relation itself changes: wear alters the vibration signature of a fault, so the same spectrum now means something else, and no amount of input monitoring alone will reveal it. The drift of a deployed model, sometimes called model drift, is the sum of both.

The usual measure of a moved input is the population stability index, computed on one variable cut into bins:

PSI=∑k (pk−qk) ln⁡pkqk,\mathrm{PSI}=\sum_k\,(p_k-q_k)\,\ln\frac{p_k}{q_k},PSI=k∑​(pk​−qk​)lnqk​pk​​,

where qkq_kqk​ is the fraction of training data in bin kkk and pkp_kpk​ the fraction of recent data. It is zero when nothing has changed; by convention below 0.1 is stable and above 0.25 is a serious shift.

Drift is watched on inputs because the labels come late. In predictive maintenance you learn whether the model was right only when the machine fails, or does not, weeks later (delayed labels, one of the traits of industrial labels); meanwhile you watch input distributions and the residuals of the physical models around it.

A drift alarm is a reason to look, and settles nothing by itself. A shifted temperature distribution in January may be harmless if the model saw last January; a stable distribution can hide concept drift. The decision to retrain goes through the MLOps loop: collect, relabel, validate, roll out in stages.

It is easy to confuse with two neighbours. An out-of-distribution input is one sample far from training, judged at that moment; drift is a property of the population over weeks. The distributional shift of a language model is a third thing again: the distribution of the next token moved by the context in the prompt.