systems theory//engineering patterns//linearize-solve-repeat

Linearize-solve-repeat is the engineering pattern by which almost every nonlinear problem in practice is solved: approximate it locally by a linear one with the Jacobian, solve that linear problem exactly, move to the answer, and approximate again there. It is used to recognise that an extended Kalman filter, Newton's method, nonlinear MPC, gain scheduling and the training of a neural network share one skeleton, and therefore one way of failing: the step leaves the region where the straight line was a good description of the curve.


Linearize-solve-repeat is the engineering pattern by which almost every nonlinear problem in practice is solved: approximate it locally by a linear one with the Jacobian, solve that linear problem exactly, move to the answer, and approximate again there. It is used to recognise that an extended Kalman filter, Newton's method, nonlinear MPC, gain scheduling and the training of a neural network share one skeleton, and therefore one way of failing: the step leaves the region where the straight line was a good description of the curve.

The root is the first-order Taylor polynomial. Near a point x0x_0x0​ any smooth function looks like its tangent.

f(x)≈f(x0)+J(x0) (x−x0)f(x) \approx f(x_0) + J(x_0)\,(x-x_0)f(x)≈f(x0​)+J(x0​)(x−x0​)

Linear problems have closed-form or very reliable solutions: a linear system of equations, a quadratic cost with a linear model, a Gaussian through a matrix. So the nonlinear problem is replaced by a sequence of linear ones, each valid near the current guess, and the sequence walks toward the answer. A drone's attitude near hover, where sin⁡θ≈θ\sin\theta\approx\thetasinθ≈θ, is the case everyone meets first (linearization).

Each method in the family is distinguished by where it relinearizes, and the answer tells you how it fails. The extended Kalman filter relinearizes at every time step around its current estimate, so a bad estimate gives a bad Jacobian and the filter can diverge. Newton's method relinearizes at every iteration, converging very fast near the root and wandering far from it. Nonlinear MPC relinearizes around the trajectory it planned on the previous cycle. Gain scheduling relinearizes offline, once per operating point, and interpolates between the designs.

Backpropagation is the pattern applied to a chain: the gradient of a deep network is a product of local linearizations, one Jacobian per layer, and each gradient descent step trusts that product only for a short distance, which is what the learning rate encodes.

Mapping and calibration problems are nonlinear least squares solved the same way. Graph SLAM linearizes every measurement constraint around the current map, solves one large sparse linear system and repeats, the Gauss-Newton loop of nonlinear least squares.

The error budget uses one pass of it without repeating: carrying a source's spread to the output with σy≈∣∂f/∂x∣ σx\sigma_y\approx|\partial f/\partial x|,\sigma_xσy​≈∣∂f/∂x∣σx​, or JΣJ⊤J\Sigma J^\topJΣJ⊤ for several, is the same linearized propagation the EKF performs at each prediction (uncertainty propagation).

The limit is local validity. Where the curvature is strong compared with the step or the uncertainty (a wide covariance through a sharply nonlinear sensor, a drone far from hover, a network in a sharp region of its loss), the tangent lies; the remedies are smaller steps, sampling instead of linearizing (unscented Kalman filter, particle filter) or a trust region. The pattern's siblings are in engineering patterns.