mathematics//linear algebra//condition number
The condition number of a matrix is the ratio of its largest to its smallest stretch, and it measures how much solving \(Ax=b\) can amplify a relative error in the data; an engineer uses it to know, before trusting a result, how many significant digits a calibration, a fit or a filter will lose. In terms of the singular values,
The condition number of a matrix is the ratio of its largest to its smallest stretch, and it measures how much solving Ax=bAx=bAx=b can amplify a relative error in the data; an engineer uses it to know, before trusting a result, how many significant digits a calibration, a fit or a filter will lose. In terms of the singular values,
κ(A)=σmaxσmin,∥δx∥∥x∥≤κ(A) ∥δb∥∥b∥.\kappa(A)=\frac{\sigma_{\max}}{\sigma_{\min}},\qquad \frac{\|\delta x\|}{\|x\|}\le\kappa(A)\,\frac{\|\delta b\|}{\|b\|}.κ(A)=σminσmax,∥x∥∥δx∥≤κ(A)∥b∥∥δb∥.
The inequality says the relative error of the answer can be up to κ\kappaκ times the relative error of the input. The rule of thumb that follows is the one to remember: you lose about log10κ\log_{10}\kappalog10κ significant digits. Double precision carries about 16 and single precision about 7 (floating-point arithmetic), so κ=106\kappa=10^6κ=106 on a microcontroller computing in float32 leaves, with luck, one reliable digit. A matrix is ill-conditioned when κ\kappaκ is large enough for rounding or sensor noise to start mattering at the precision in use. A rotation has κ=1\kappa=1κ=1, the best possible; a matrix that has lost rank has κ=∞\kappa=\inftyκ=∞.
Scale first, never invert, and use solve or lstsq.
Fitting a degree-5 polynomial on raw inputs from 0 to 100 gives a matrix with κ≈1.6×1010\kappa\approx1.6\times10^{10}κ≈1.6×1010, ten of sixteen digits gone; rescaling the inputs to [−1,1][-1,1][−1,1] brings it to about 41. Half of numerical problems are fixed by variable scaling before any algorithm is touched.
Forming ATAA^{\mathsf T}AATA squares the condition number, so the textbook normal equations of least squares lose twice the digits; the same polynomial gives κ(ATA)≈2.7×1020\kappa(A^{\mathsf T}A)\approx2.7\times10^{20}κ(ATA)≈2.7×1020, which in double precision is garbage. QR and the SVD work on AAA itself.
It is the how much behind several yes-or-no tests. The smallest singular value of the observability matrix says how badly the worst direction is seen (observability); a magnetometer calibrated only in level flight has its vertical axis barely excited and comes out with a noisy correction (magnetometer). An ill-conditioned calibration matrix amplifies sensor noise, which reaches the Kalman filter as a wrong RRR and from there the controller.
The same ratio of curvatures, taken on a Hessian, sets how badly gradient descent zigzags down a long narrow valley; standardizing features brings it towards 1.
Conditioning belongs to the problem, stability to the algorithm. An ill-conditioned problem loses digits whatever the method; an unstable method loses them even on a well-conditioned problem, and forming an explicit inverse is the usual way to add the second loss to the first (numerical stability).