ML//generalization//regularization//lasso
Lasso is linear regression with a penalty on the sum of the absolute values of the weights, and its signature effect is that many weights land exactly at zero, so it selects features while it fits; it is used when a model has more candidate inputs than the data can support (dozens of sensor channels and lags for a few hundred labelled cycles) and the engineer wants to know which few matter. Tibshirani named it in 1996, for *least absolute shrinkage and selection operator*.
Lasso is linear regression with a penalty on the sum of the absolute values of the weights, and its signature effect is that many weights land exactly at zero, so it selects features while it fits; it is used when a model has more candidate inputs than the data can support (dozens of sensor channels and lags for a few hundred labelled cycles) and the engineer wants to know which few matter. Tibshirani named it in 1996, for least absolute shrinkage and selection operator.
minw ∑i=1N(yi−w⊤xi)2+λ∥w∥1,∥w∥1=∑j∣wj∣\min_{w}\;\sum_{i=1}^{N}\big(y_i-w^{\top}x_i\big)^2+\lambda\lVert w\rVert_1,\qquad \lVert w\rVert_1=\sum_j\lvert w_j\rvertwmini=1∑N(yi−w⊤xi)2+λ∥w∥1,∥w∥1=j∑∣wj∣
Here λ\lambdaλ is the price of complexity: zero gives ordinary least squares, and as it grows the weights are switched off one by one. The zeros come from the shape of the penalty. An absolute value has a corner at zero, so a weight whose contribution to the fit is worth less than λ\lambdaλ finds its cheapest place exactly at zero; the squared penalty of ridge regression is smooth there and only shrinks the weight toward it.
Its output is a sparse model that reads like a short list: of forty process variables, the six that predict the energy of a heat-treatment furnace, each weight with its units. That list is auditable and cheap to run, a dot product over six inputs on any controller.
Among strongly correlated features it keeps one, somewhat arbitrarily, and drops the rest; with plant sensors tied by mass and energy balances (redundant sensors) the chosen one can change from one data sample to the next. The elastic net, which adds a ridge term, keeps groups of correlated features together.
A zero weight means the feature was not needed given the others in these data, which leaves open whether it matters physically. Inputs are standardized first, since the penalty treats every weight on the same scale (variable scaling), and λ\lambdaλ is chosen by cross-validation (regularization).