ML//generalization//bias-variance trade-off
The bias-variance trade-off is the decomposition of a model's expected squared error into a part due to the model being too rigid (bias), a part due to its sensitivity to the particular sample it was trained on (variance), and a noise floor no model goes below, and it is what guides the choice of model complexity for a given amount of data. A model with high bias is wrong the same way on any dataset, like a stopped clock; that is **underfitting**, a straight line through a curve. A model with high variance fits the noise and changes its mind with every new sample, like a weather vane.
The bias-variance trade-off is the decomposition of a model's expected squared error into a part due to the model being too rigid (bias), a part due to its sensitivity to the particular sample it was trained on (variance), and a noise floor no model goes below, and it is what guides the choice of model complexity for a given amount of data. A model with high bias is wrong the same way on any dataset, like a stopped clock; that is underfitting, a straight line through a curve. A model with high variance fits the noise and changes its mind with every new sample, like a weather vane.
At a point xxx, averaging over all the training sets that could have been drawn,
E[(y−f^(x))2]=(f(x)−E[f^(x)])2⏟bias2+Var[f^(x)]⏟variance+σ2\mathbb E\big[(y-\hat f(x))^2\big]=\underbrace{\big(f(x)-\mathbb E[\hat f(x)]\big)^2}_{\text{bias}^2}+\underbrace{\operatorname{Var}\big[\hat f(x)\big]}_{\text{variance}}+\sigma^2E[(y−f^(x))2]=bias2(f(x)−E[f^(x)])2+varianceVar[f^(x)]+σ2
where fff is the true relation, f^\hat ff^ the fitted model (random, because it depends on which data came in) and σ2\sigma^2σ2 the irreducible error, the noise of the measurement itself. More complexity lowers bias and raises variance; more data lowers variance. Hence the U-shaped curve of test error against complexity.
degree 3: training0.186 test0.280 best degree3 parameters / points4 / 16 Polynomials fitted to 16 noisy samples of a sine. Training error falls with every degree, from 0.597 at degree 0 to 0.000 at degree 15; test error bottoms out at degree 3 (0.280, near the noise floor of 0.25) and reaches 289 at degree 15.
Raise the degree from 0 to 15: the training error always falls, the test error bottoms out and then shoots up. At a high degree press New data a few times and watch the curve dance, which is the variance; then raise λ\lambdaλ, then the number of points, and follow the best degree.
The noise floor says when to stop. When test error approaches σ2\sigma^2σ2, the model is done and further gains belong to the sensor; repeated measurements at one fixed operating point estimate that floor before any model is fitted (variance).
With few data, the cheapest variance reduction is structure. A degree-12 polynomial through 15 measurements of a pump's efficiency passes through every point and predicts nonsense between them, while a degree-2 curve with the expected physical shape is right (grey-box model). The knobs that move along the trade-off are regularization, early stopping and pruning.
Double descent is where the intuition fails. Overparametrized networks generalize better than the U-curve predicts, with test error falling again past the point where the model can fit its training data exactly. The decomposition still holds; what fails is the reading more parameters, more variance (overfitting).