mathematics//linear algebra//least squares

Least squares is the method that fits the parameters of a linear model to more measurements than parameters by minimizing the sum of squared residuals, and it is the workhorse behind sensor calibration, linear regression, curve fitting and every recursive estimator built on it. With more equations than unknowns, \(Ax=b\) has no exact solution: \(b\) lies outside the **column space** of \(A\), the set of all outputs \(Ax\) the model can produce. The best available answer is the point of that space closest to \(b\), and the closest point of a plane is the orthogonal projection: the residual \(r=b-A\hat x\) must be perpendicular to every column.


Least squares is the method that fits the parameters of a linear model to more measurements than parameters by minimizing the sum of squared residuals, and it is the workhorse behind sensor calibration, linear regression, curve fitting and every recursive estimator built on it. With more equations than unknowns, Ax=bAx=bAx=b has no exact solution: bbb lies outside the column space of AAA, the set of all outputs AxAxAx the model can produce. The best available answer is the point of that space closest to bbb, and the closest point of a plane is the orthogonal projection: the residual r=b−Ax^r=b-A\hat xr=b−Ax^ must be perpendicular to every column.

AT(b−Ax^)=0⟺ATA x^=ATbA^{\mathsf T}(b-A\hat x)=0 \quad\Longleftrightarrow\quad A^{\mathsf T}A\,\hat x=A^{\mathsf T}bAT(b−Ax^)=0⟺ATAx^=ATb

This is why they are called the normal equations: they say the residual is normal to the column space. A calibration of an accelerometer works exactly so. Each reading taken at rest in a known orientation is one row, the unknowns are scale factors, misalignments and offsets, and a few dozen orientations overdetermine a dozen parameters.

Least squares is an orthogonal projection, and regularizing it switches off the directions the data do not see.

The geometry is the point; the formula is better left unused. Forming ATAA^{\mathsf T}AATA squares the condition number, so the solve goes through QR or the SVD (np.linalg.lstsq).

The square carries a statistical assumption. Minimizing the squared residual is maximum likelihood when the noise is Gaussian, independent and of constant variance. When each point has its own variance, weighted least squares divides each residual by its standard deviation, the same weights as inverse-variance weighting in sensor fusion.

Directions the data barely see are where the noise goes. Through the SVD the solution divides the data's component along each direction by its singular value, so a nearly invisible direction multiplies its noise without mercy; ridge regression damps those directions and persistent excitation is the experimental fix that makes them visible in the first place.

Polynomial fitting is the classic trap: the matrix of powers of xxx is notoriously ill-conditioned on unscaled data, and a fit that looks fine in double precision collapses in single (variable scaling).

The square inherits a weakness for outliers: one wild point pulls the projection hard, because its squared residual dominates the sum. With outliers in the data, the losses of robust statistics or RANSAC take over.

The method has relatives for each situation it does not cover: recursive least squares updates the fit sample by sample, nonlinear least squares linearizes and repeats when the model is not linear in its parameters.