mathematics//linear algebra//pseudoinverse

The pseudoinverse of a matrix \(A\), written \(A^+\), is the generalization of the inverse to matrices that are not square or not invertible, defined so that \(x=A^+b\) is the best available answer to \(Ax=b\): the least squares fit when there are more equations than unknowns, and the smallest solution when there are fewer. Engineers meet it whenever a linear problem has the wrong shape: fitting a calibration to many readings, splitting four control demands among six motors, inverting a robot arm's Jacobian.


The pseudoinverse of a matrix AAA, written A+A^+A+, is the generalization of the inverse to matrices that are not square or not invertible, defined so that x=A+bx=A^+bx=A+b is the best available answer to Ax=bAx=bAx=b: the least squares fit when there are more equations than unknowns, and the smallest solution when there are fewer. Engineers meet it whenever a linear problem has the wrong shape: fitting a calibration to many readings, splitting four control demands among six motors, inverting a robot arm's Jacobian.

The two cases are the two ways a linear system of equations can fail to have one exact answer. With more independent rows than columns (a dozen calibration parameters, sixty readings), no xxx satisfies every equation, and A+bA^+bA+b minimizes the squared residual; with full column rank it equals (ATA)−1ATb(A^{\mathsf T}A)^{-1}A^{\mathsf T}b(ATA)−1ATb. With more columns than independent rows, infinitely many xxx satisfy the equations (any vector of the null space can be added), and A+bA^+bA+b picks the one of smallest length, AT(AAT)−1bA^{\mathsf T}(AA^{\mathsf T})^{-1}bAT(AAT)−1b when the rows are independent. A hexacopter's mixer is the second case: six motor thrusts produce four quantities (total thrust and three torques), and the pseudoinverse of the 4×6 allocation matrix gives the motor commands that achieve the demand with the least total effort (control allocation).

It is computed through the SVD, and its danger is the same as the SVD's.

A+=VΣ+UTA^+=V\Sigma^+U^{\mathsf T}A+=VΣ+UT, where Σ+\Sigma^+Σ+ inverts every nonzero singular value (SVD). A singular value that is tiny but nonzero is inverted into a huge one, so a nearly dependent pair of columns turns measurement noise into wild parameters; practical code (np.linalg.pinv with its rcond) treats singular values below a tolerance as zero, which is regularization by truncation.

The formulas with (ATA)−1(A^{\mathsf T}A)^{-1}(ATA)−1 explain the result and are a poor way to compute it: forming ATAA^{\mathsf T}AATA squares the condition number, and an explicit pseudoinverse is rarely needed when a solve will do (np.linalg.lstsq returns the same minimum-norm solution more cheaply).

Minimum norm is a choice, and not always the right one. In control allocation it ignores motor limits, so after saturation the pseudoinverse answer must be clipped and redistributed, and modern allocators solve a small constrained optimization instead; in a robot arm near a singular pose it demands huge joint speeds, and a damped version (adding λ2I\lambda^2 Iλ2I before inverting, the same move as ridge regression) trades a little accuracy for bounded commands.

A matrix with full rank and square shape has A+=A−1A^+=A^{-1}A+=A−1, so the pseudoinverse never gives a different answer where the ordinary inverse exists.