mathematics//linear algebra//matrix factorization//SVD
The SVD (singular value decomposition) is the factorization of any matrix into a rotation, a stretch along the axes and another rotation, and it is the tool that says how much a matrix can amplify, which directions it barely sees and how many it really keeps; eigenvalues cannot answer those questions, and they exist only for square matrices. For any \(m\times n\) matrix,
The SVD (singular value decomposition) is the factorization of any matrix into a rotation, a stretch along the axes and another rotation, and it is the tool that says how much a matrix can amplify, which directions it barely sees and how many it really keeps; eigenvalues cannot answer those questions, and they exist only for square matrices. For any m×nm\times nm×n matrix,
A=UΣVT,σ1≥σ2≥⋯≥0.A=U\Sigma V^{\mathsf T},\qquad \sigma_1\ge\sigma_2\ge\dots\ge0 .A=UΣVT,σ1≥σ2≥⋯≥0.
VVV and UUU are orthogonal matrices (they turn, perhaps with a reflection, without deforming) and Σ\SigmaΣ is diagonal with the singular values σi\sigma_iσi. Geometrically, AAA turns the unit sphere into an ellipsoid whose semi-axes measure σi\sigma_iσi. The number of nonzero singular values is the rank; the number above a sensible tolerance is the numerical rank, the one that counts with data; the ratio of the largest to the smallest is the condition number. Everything a matrix does, however ugly, is read off one diagonal.
The SVD measures how much a matrix can amplify, which its eigenvalues do not tell.
A stable system with all eigenvalues at 0.9 can still multiply a disturbance by 19 on the way to rest (non-normal matrix); the largest singular value of AkA^kAk shows it, the eigenvalues never will.
In least squares it shows where the danger lives. The solution is x^=∑iuiTbσivi\hat x=\sum_i \frac{u_i^{\mathsf T}b}{\sigma_i}v_ix^=∑iσiuiTbvi: the data's component along each direction is divided by its singular value, so a direction the data barely see multiplies its noise without mercy. Truncated SVD drops the smallest singular values outright, regularization with scissors; ridge regression damps them smoothly. The same sum with only the nonzero terms is the pseudoinverse, which gives the minimum-norm solution of an underdetermined system.
Its uses all follow from that diagonal. PCA is the SVD of the centred data matrix, preferred to the eigendecomposition of the covariance because it avoids squaring the conditioning. The smallest singular value of the observability matrix says how badly the worst direction is seen (observability), that of the controllability matrix how much effort the hardest direction costs. In a deep network the gradient passes through a product of layer Jacobians, and singular values consistently below or above 1 make it vanish or explode (vanishing gradient).
It is the expensive member of the matrix factorization family: a dense 10,000 × 1,000 SVD takes of the order of seconds on a laptop, and on a microcontroller there is usually no SVD at all. For huge sparse matrices only the few largest singular values are computed, with iterative methods (scipy.sparse.linalg.svds).