ML//MLOps//training-serving skew
Training-serving skew is the difference between the model that was trained and validated and the model that actually runs in production, caused by anything in the path from raw signal to prediction that is not identical on both sides, and checking for it is the last step before a model is trusted in the field. The book calls it the silliest and most frequent error of deployment, because each cause is mundane and each one silently invalidates the test score.
Training-serving skew is the difference between the model that was trained and validated and the model that actually runs in production, caused by anything in the path from raw signal to prediction that is not identical on both sides, and checking for it is the last step before a model is trusted in the field. The book calls it the silliest and most frequent error of deployment, because each cause is mundane and each one silently invalidates the test score.
Three causes account for most cases. The preprocessing drifts apart: production computes features with different units, a different column order, a different filter or a different sampling rate than the training notebook did (a detector developed with zero-phase filtering offline arrives late once deployed with a causal filter). The artifact changes: a network quantized to int8 for the edge is a different model with its own errors, and so is one exported to ONNX or to generated C. The sensor changes: the prototype's accelerometer is not the series one, with another bias, noise floor and mounting.
Validate the artifact you deploy, with the preprocessing you deploy, on data from the machine where you deploy it. Any validation run on something else measured something else.
The cheapest defence is shared code: one feature library used by training and serving, versioned with the model (the rule of MLOps that every behaviour-changing piece is versioned), plus a check that recomputes a batch of production features offline and compares them number by number.
It is distinct from data drift. Drift is the world moving after deployment; skew is a mismatch present from the first second, which is why a model can underperform on day one with no drift at all.
Shadow mode catches skew only if the shadow runs the deployed artifact on the deployed hardware; a shadow running the training model on a server would reproduce the validation score and hide the problem.