mathematics//numerical methods//variable scaling

Variable scaling is the change of units applied to the inputs of a computation (shifting and stretching each variable so that all of them live in similar ranges, usually around \([-1,1]\) or with zero mean and unit variance) so that the numbers the algorithm handles are well balanced, and it is the cheapest fix in numerical work: about half of numerical problems disappear when the variables are scaled, before any algorithm is changed. It alters nothing physical; the model, the data and the answer are the same, expressed in better units.


Variable scaling is the change of units applied to the inputs of a computation (shifting and stretching each variable so that all of them live in similar ranges, usually around [−1,1][-1,1][−1,1] or with zero mean and unit variance) so that the numbers the algorithm handles are well balanced, and it is the cheapest fix in numerical work: about half of numerical problems disappear when the variables are scaled, before any algorithm is changed. It alters nothing physical; the model, the data and the answer are the same, expressed in better units.

A polynomial fit shows the size of the effect. Fitting a degree-5 polynomial to times between 0 and 100 s builds a Vandermonde matrix whose columns range from 1 to 1005=1010100^5=10^{10}1005=1010; its condition number is about 1.6×10101.6\times10^{10}1.6×1010, so roughly ten of the sixteen digits of double precision are lost before the fit starts. Rescaling time to [−1,1][-1,1][−1,1] with xs=(x−50)/50x_s=(x-50)/50xs​=(x−50)/50 brings the condition number down to 41, and the same fit loses less than two digits. On a microcontroller in single precision the unscaled version returns noise.

Feature standardization is the same move in machine learning: subtract each feature's mean and divide by its standard deviation. It brings the condition number of the problem towards one, which matters twice. Gradient descent on badly scaled features crawls along a long narrow valley, because the steepest direction caps the step and the flattest sets the pace; and methods built on distances (k-nearest neighbors, support vector machines, PCA) are dominated by whichever variable has the largest units, so a pressure in pascals beside a temperature in degrees decides everything by itself.

The statistics used for scaling must come from the training data only and travel with the model. Scaling a test set with its own mean leaks information, and a model deployed without the training means and deviations receives inputs in the wrong units, a quiet form of training-serving skew (training-serving skew).

Some methods do not need it. Decision trees and gradient boosting only compare each feature with thresholds, so a monotone change of units leaves them unchanged; that is one reason they are forgiving on raw plant data.

In physics and control the same idea is nondimensionalization: expressing states in units of their nominal value or their range (per-unit quantities in power systems, angles in radians, positions in metres instead of millimetres). In fixed-point arithmetic scaling becomes part of the design of every variable.