ML//supervised learning//linear regression
Linear regression is a supervised model that predicts a number as a weighted sum of features, with the weights chosen to minimize the squared error on the examples; it is used to calibrate sensors, fit consumption curves and, in every machine learning project, as the auditable baseline any other model must beat. The energy of a heat-treatment furnace is predicted from the load, the setpoint and the ambient temperature:
Linear regression is a supervised model that predicts a number as a weighted sum of features, with the weights chosen to minimize the squared error on the examples; it is used to calibrate sensors, fit consumption curves and, in every machine learning project, as the auditable baseline any other model must beat. The energy of a heat-treatment furnace is predicted from the load, the setpoint and the ambient temperature:
y^=wTx+b=w1x1+⋯+wdxd+b.\hat y=w^{\mathsf T}x+b=w_1x_1+\cdots+w_dx_d+b .y^=wTx+b=w1x1+⋯+wdxd+b.
xxx holds the ddd features, www the weights, bbb the intercept and y^\hat yy^ the prediction. Each weight carries units (kWh per kilogram of load, kWh per degree of setpoint) and can be checked by eye: if the load weight comes out negative, the problem is in the data. Minimizing the squared error is the least squares problem, and under Gaussian noise it is also the maximum likelihood estimate. It is solved with QR or the SVD, never by inverting XTXX^{\mathsf T}XXTX, which squares the condition number and throws away digits.
Linear refers to the parameters, not the variables.
Squares, products and physical transforms of the inputs (x2x^2x2, ω2\omega^2ω2, I2I^2I2 for Joule losses) can be added as new columns and the model is still linear regression, solved the same way. Writing a model as measured regressors times unknown parameters is what lets least squares and recursive least squares identify a drone's thrust coefficients or a motor's resistance online.
It fails with interactions nobody wrote. If the effect of load depends on the setpoint, the model averages it away unless a product column is added; trees capture such interactions on their own, which is part of why they win on rich tables.
Squared loss gives outliers outsized weight. One reading at ten times the usual residual pulls the fit a hundred times harder than a typical one; with spikes in the data, a robust loss from robust statistics keeps the fit on the bulk.
Multicollinearity blows up the weights. Near-collinear features (two thermocouples a few centimetres apart, or ω\omegaω and ω2\omega^2ω2 over a narrow speed range) make the condition number explode, and the weights come out large, of opposite signs and unstable from one sample to the next. The predictions may still be fine; the interpretation is not. More informative data fix it (experiment design), and a penalty hides it (ridge regression).
It costs one dot product to run, which fits a one-euro microcontroller at 1 kHz, and it trains in seconds on as little as ten or twenty examples per feature (a common rule of thumb for stable weights).
Its classifier sibling is logistic regression; its place in the family is set out in supervised learning.