mathematics//linear algebra//linear system of equations

A linear system of equations is a set of equations in which the unknowns appear only multiplied by constants and added, written \(Ax=b\), and it is the form that calibration, estimation, circuit analysis and every step of a nonlinear solver end up in. The question it asks is which input \(x\) the transformation \(A\) turns into the observed \(b\), and the answer depends on the shape of \(A\) far more than on its numbers.


A linear system of equations is a set of equations in which the unknowns appear only multiplied by constants and added, written Ax=bAx=bAx=b, and it is the form that calibration, estimation, circuit analysis and every step of a nonlinear solver end up in. The question it asks is which input xxx the transformation AAA turns into the observed bbb, and the answer depends on the shape of AAA far more than on its numbers.

Three cases cover what happens, and none needs memorizing once the matrix is read as a machine. If AAA is square and crushes nothing (full rank), every bbb comes from exactly one xxx. If there are more equations than unknowns, which is the normal situation with real data (two hundred measurements for two parameters), bbb almost never lies among the outputs AAA can produce, so there is no exact solution and the best one can do is the closest output: that is least squares. If there are fewer equations than unknowns, any vector of the null space can be added to a solution without changing AxAxAx, so there are infinitely many, and a criterion must pick one. The usual pick is the minimum-norm solution, the one of smallest length, which the pseudoinverse computes through the SVD; a robot arm with more joints than the task needs is solved this way, the null space being the motions of the elbow that leave the hand where it is.

Never invert, solve.

The call np.linalg.solve factorizes AAA (LU decomposition) and substitutes, which costs less than forming A−1A^{-1}A−1 and multiplying and is numerically more stable. The explicit inverse adds nothing and amplifies rounding error on the way.

How much an exact solution can be trusted is a separate question from whether it exists. A square system can be solvable on paper and useless in practice when the condition number of AAA is large: the relative error of xxx can be that many times the relative error of bbb, so a noisy sensor reading turns into a wild estimate.

The same three cases come back under other names. An unobservable state is a null-space direction of the observability matrix (observability); an overdetermined fit with nearly dependent columns is an identifiability problem that more computation does not fix (persistent excitation); a Newton step inside nonlinear least squares is one linear solve per iteration.

Which factorization runs under solve depends on the matrix: LU for a general square one, Cholesky for a symmetric positive definite one such as a covariance, QR for least squares. The family is gathered in matrix factorization.