ML//generalization

Generalization is a model's performance on data it did not see during training, drawn from the conditions where it will be used, and it is the real objective of machine learning although it appears nowhere in the formula the model is trained on. Empirical risk minimization minimizes the average loss on the examples at hand; what an engineer wants is a model that is right next winter, on the next machine, in the next plant.


Generalization is a model's performance on data it did not see during training, drawn from the conditions where it will be used, and it is the real objective of machine learning although it appears nowhere in the formula the model is trained on. Empirical risk minimization minimizes the average loss on the examples at hand; what an engineer wants is a model that is right next winter, on the next machine, in the next plant.

The gap between the two is the generalization error, and its sources look alike from outside. A model can be too rigid or too flexible for the amount of data it has (bias-variance trade-off); too flexible, it fits the noise, which is overfitting, and the remedy is a price on complexity (regularization, whose two classic forms are ridge regression and lasso) or more data. A model can be well fitted and meet inputs unlike any it saw (out-of-distribution). Or it can have learned the wrong regularity, the easiest one that separated its training classes (shortcut learning). The first is cured by a better fit; the other two show up only under the right test (train-validation-test split).

A learned model that replaces physics moves the modelling error instead of removing it.

The error goes down on the test data and reappears as generalization error the day the drone carries a heavier parcel than any in its training set (error relocation).

Physics is the cheapest generalization there is. A grey-box model keeps working, degraded, outside its data because the physics still holds, while a black box can do anything there (grey-box model). Features that already discount known dependencies give a learned model part of that invariance (physics-based features).

Coverage beats volume. A thousand summer hours say nothing about winter, and a historian full of correlated samples holds far less information than its row count suggests (dataset, effective sample size).

Generalization is measured, never assumed: honest evaluation estimates it before deployment, and data drift monitoring watches it erode afterwards.