ML//generalization//regularization//ridge regression

Ridge regression is linear least squares with a penalty on the sum of the squared weights, and it is the standard cure for regressions whose inputs are nearly collinear or too numerous for the data: it shrinks every weight toward zero and keeps the fit stable where plain least squares would hand back huge coefficients driven by noise. It is the same computation as Tikhonov regularization in inverse problems and as the MAP estimate with a Gaussian prior on the weights.


Ridge regression is linear least squares with a penalty on the sum of the squared weights, and it is the standard cure for regressions whose inputs are nearly collinear or too numerous for the data: it shrinks every weight toward zero and keeps the fit stable where plain least squares would hand back huge coefficients driven by noise. It is the same computation as Tikhonov regularization in inverse problems and as the MAP estimate with a Gaussian prior on the weights.

min⁡w  ∑i=1N(yi−w⊤xi)2+λ∥w∥22\min_{w}\;\sum_{i=1}^{N}\big(y_i-w^{\top}x_i\big)^2+\lambda\lVert w\rVert_2^2wmin​i=1∑N​(yi​−w⊤xi​)2+λ∥w∥22​

Its effect is clearest through the SVD of the data matrix. Plain least squares divides the data's component along each singular direction by its singular value σi\sigma_iσi​, so a direction the data barely see (a tiny σi\sigma_iσi​) multiplies its noise without mercy. Ridge multiplies each of those components by a filter factor,

σi2σi2+λ\frac{\sigma_i^{2}}{\sigma_i^{2}+\lambda}σi2​+λσi2​​

which is close to 1 where σi2≫λ\sigma_i^2\gg\lambdaσi2​≫λ and close to 0 where σi2≪λ\sigma_i^2\ll\lambdaσi2​≪λ: it switches off exactly the directions the data cannot determine and leaves the well-measured ones alone.

In algebraic terms it adds λ\lambdaλ to every eigenvalue of X⊤XX^{\top}XX⊤X, which bounds its condition number. Two thermocouples a centimetre apart in a furnace regression can otherwise get large weights of opposite sign that nearly cancel; ridge shares the effect between them.

The Bayesian reading sets the price: with noise variance σε2\sigma_\varepsilon^2σε2​ and a prior spread τ\tauτ on each weight, λ=σε2/τ2\lambda=\sigma_\varepsilon^2/\tau^2λ=σε2​/τ2. A datasheet that promises a gain of 1 ± 0.05 is such a prior, and with it four noisy calibration points can no longer drag the gain to 0.83 (regularization).

It never sets a weight exactly to zero, so every input stays in the model; when a short list of features is wanted, lasso gives it. Dropping the small singular values outright, the truncated SVD, is the version with scissors. Inputs are standardized first, and λ\lambdaλ is chosen by cross-validation.