ML//generalization//shortcut learning
Shortcut learning is the failure in which a model learns the easiest regularity that separates its training classes instead of the one its designers meant, and it is what makes a model that scored well in validation fail on its first day in a new setting. If every photo of a defective part was taken on the same day under different lighting, the inspection network learns to detect the lighting. Geirhos and colleagues gave the pattern its name in 2020.
Shortcut learning is the failure in which a model learns the easiest regularity that separates its training classes instead of the one its designers meant, and it is what makes a model that scored well in validation fail on its first day in a new setting. If every photo of a defective part was taken on the same day under different lighting, the inspection network learns to detect the lighting. Geirhos and colleagues gave the pattern its name in 2020.
Nothing in training tells the intended cue from the accidental one: both lower the loss, and the accidental one is usually simpler (a brightness level, a watermark, a timestamp, the hospital a scan came from, a sensor channel fitted only on the machines that later failed). Empirical risk minimization rewards whichever separates the examples, and a random split cannot expose it, because the shortcut is present on both sides of the split. The model is right for the wrong reason, and its score says nothing about the reason.
It is exposed by group validation: test on another day, another line, another camera, another plant, anything that breaks the accidental correlation while keeping the real one (train-validation-test split). A score that collapses across groups is the signature.
It is also exposed by looking at what the model uses: saliency maps over pixels for images, attributions over time intervals for signals, SHAP values for tables. A defect classifier whose attention sits on the background has found a shortcut.
In tables its cousin is label leakage, a feature that is a consequence of the answer (data leakage). The two differ in where they live: a leak lives in the data pipeline, a shortcut in the world that produced the data, and it survives a perfectly clean pipeline.
The remedy lies in the data more than in the model: collect examples where the shortcut and the label disagree, randomize the nuisance factor during collection, or augment it away (data augmentation). Shortcuts are one reason a network answers confidently and wrongly once it leaves its distribution (out-of-distribution).