mathematics//calculus//chain rule
The chain rule is the rule for differentiating a composition of functions: when \(z\) depends on \(y\) and \(y\) depends on \(x\), the rate of \(z\) with respect to \(x\) is the product of the two intermediate rates. Engineers use it whenever an effect travels through a chain of stages, to know how much the end moves when the start is nudged, and it is the whole mathematical content of training a neural network.
The chain rule is the rule for differentiating a composition of functions: when zzz depends on yyy and yyy depends on xxx, the rate of zzz with respect to xxx is the product of the two intermediate rates. Engineers use it whenever an effect travels through a chain of stages, to know how much the end moves when the start is nudged, and it is the whole mathematical content of training a neural network.
dzdx=dzdy dydx.\frac{dz}{dx}=\frac{dz}{dy}\,\frac{dy}{dx}.dxdz=dydzdxdy.
A motor's speed sets a pump's flow and the flow sets a tank's level rate. If the flow rises 0.2 L/s per 100 rpm and the level rate rises 0.4 mm/s per L/s, the level rate rises 0.08 mm/s per 100 rpm: the gains of the stages multiply. With several variables each factor becomes a Jacobian and the product a matrix product, Jz,x=Jz,y Jy,xJ_{z,x}=J_{z,y},J_{y,x}Jz,x=Jz,yJy,x; when a quantity reaches the output by several paths, the contributions of the paths add.
Backpropagation is the chain rule applied in the right order with the intermediate results written down. A network is a composition of layers; the derivative of the loss with respect to a weight deep inside is the product of the derivatives of every layer after it. Multiplying from the loss backwards shares the partial products among all the weights, which is why the gradient for a billion parameters costs only a small multiple of one forward pass. It is a way of computing derivatives, and the learning is done by gradient descent afterwards.
A long product is a stability question. Multiplying fifty factors slightly below one sends the result to zero, and fifty slightly above one sends it to infinity; that is the vanishing gradient of deep and recurrent networks, read as the stability of a discrete system. Residual connections add an identity to each factor so the product cannot collapse.
The same rule carries uncertainty through a measurement chain. A strain gauge's resistance noise passes through a bridge, an amplifier and a calibration before it becomes a force; to first order the noise at the end is the noise at the start times the product of the stage sensitivities (uncertainty propagation).
Done by machine, the chain rule is automatic differentiation: the program records each elementary operation and applies the rule to every one of them, forwards or backwards.