mathematics//linear algebra//linear transformation

A linear transformation is a function between vector spaces that preserves sums and scalings, and a matrix is how it is written down once coordinates are chosen; reading a matrix this way, as a machine that moves vectors, is what lets an engineer guess what a calibration, a dynamics matrix or a network layer does before computing anything. Coming from spreadsheets, a matrix looks like a table of numbers. The useful reading is by columns: \(Ax\) mixes the columns of \(A\) with the weights given by \(x\).


A linear transformation is a function between vector spaces that preserves sums and scalings, and a matrix is how it is written down once coordinates are chosen; reading a matrix this way, as a machine that moves vectors, is what lets an engineer guess what a calibration, a dynamics matrix or a network layer does before computing anything. Coming from spreadsheets, a matrix looks like a table of numbers. The useful reading is by columns: AxAxAx mixes the columns of AAA with the weights given by xxx.

Ax=x1a1+x2a2+⋯+xnanAx = x_1 a_1 + x_2 a_2 + \dots + x_n a_nAx=x1​a1​+x2​a2​+⋯+xn​an​

Here aja_jaj​ is column jjj of AAA. Feed in the unit vector along axis jjj and what comes out is aja_jaj​: column jjj is where axis jjj lands. With that reading a matrix can be classified by what it does to space. A scaling matrix is diagonal and stretches each axis by its own factor; a rotation turns space without deforming it; a projection flattens it onto a plane; the dynamics matrix of a state-space model, xk+1=Axkx_{k+1}=Ax_kxk+1​=Axk​, is one time step written as a transformation.

Read a matrix as a machine that moves vectors.

Its columns say where each axis goes, its null space says what it destroys beyond recovery, and its rank says how many independent directions survive the trip.

Linear is the approximation that can be solved, which is why it is everywhere even though the world is not linear. Near an operating point any smooth system behaves like its Jacobian, which is a matrix (linearization). The calibration of a three-axis accelerometer (scale factors on the diagonal, misalignment between axes off it) is a 3 × 3 matrix plus an offset vector. Each layer of a neural network is a matrix followed by a simple nonlinearity.

Composition is multiplication, and the order matters. Applying BBB and then AAA is the single matrix ABABAB, read right to left. Rotating and then stretching along the x axis is a different machine from stretching and then rotating: the first stretches along the original axis, the second along the rotated one. Code that builds a chain of frame changes (sensor to body, body to Earth) breaks in exactly this place when two factors are swapped.

A transformation and its matrix are different things. Changing the basis changes the numbers in the matrix and leaves the operation untouched; the eigendecomposition is the choice of basis in which a transformation becomes a pure scaling, and the SVD shows that every matrix is a rotation, a scaling and another rotation.

What the transformation does to volumes is its determinant; whether Ax=bAx=bAx=b can be undone is the subject of the linear system of equations.