control//adaptive control//self-tuning regulator

A self-tuning regulator is an indirect adaptive controller that estimates the plant's model parameters online and redesigns the controller at every step as if the latest estimate were exactly right, and it is the way adaptation most often reaches production: as one or two physical parameters estimated in the background and fed to an otherwise ordinary controller. Åström and Wittenmark introduced it in 1973. It separates two jobs that MRAC merges: identify the plant, then design for it. Acting on the estimate as if it were true is the certainty equivalence principle.


A self-tuning regulator is an indirect adaptive controller that estimates the plant's model parameters online and redesigns the controller at every step as if the latest estimate were exactly right, and it is the way adaptation most often reaches production: as one or two physical parameters estimated in the background and fed to an otherwise ordinary controller. Åström and Wittenmark introduced it in 1973. It separates two jobs that MRAC merges: identify the plant, then design for it. Acting on the estimate as if it were true is the certainty equivalence principle.

The estimator is usually recursive least squares: write the model as a regression yk=φkTθ+vky_k=\varphi_k^{\mathsf T}\theta+v_kyk​=φkT​θ+vk​, with φk\varphi_kφk​ the measured regressors and θ\thetaθ the parameters, and update θ^\hat\thetaθ^ at every sample with a forgetting factor λ\lambdaλ that sets a memory of about 1/(1−λ)1/(1-\lambda)1/(1−λ) samples (with λ=0.99\lambda=0.99λ=0.99 at 100 Hz, roughly the last second). RLS is a Kalman filter whose state does not move, so a self-tuning regulator is a parameter Kalman filter plus any design method for the controller (LQR, pole placement, a PID rule).

A drone shows the humblest version. Near hover the accelerometer measures a vertical specific force equal to thrust over mass, fz≈βuf_z\approx\beta ufz​≈βu, with uuu the thrust command and β=kT/m\beta=k_T/mβ=kT​/m the thrust efficiency. A scalar RLS with φ=u\varphi=uφ=u estimates β^\hat\betaβ^​, and the hover thrust becomes g/β^g/\hat\betag/β^​; that estimate corrects the feedforward and rescales the gains when a parcel is loaded or the battery sags. PX4 ships a hover-thrust estimator of this kind.

Bad initial estimates feed themselves: a poor model gives a poor controller, which produces poorer data. Start from good priors and bound the parameters (robust adaptation).

Covariance windup is its classic failure. With forgetting and no excitation PPP grows like λ−k\lambda^{-k}λ−k, and when a manoeuvre finally arrives the gain is huge and the parameters jump. The fixes are to cap the trace of PPP or to forget only in the directions the data excite (persistent excitation).

The adapted parameters are also health indicators. If one drone's estimated thrust efficiency sits 15 % below the fleet's, suspect a propeller or a motor (health indicator); the adapter already did the estimating and only the comparison is missing.

Commercial PID autotuners (and the autotune modes of ArduPilot and PX4) are its batch cousin: excite once, fit a simple model, compute gains, stop adapting (system identification).