control//state estimation//recursive least squares
Recursive least squares (RLS) is an online estimator of the parameters of a model that is linear in them, updating the least squares solution with each new sample without storing the history, and it is used to follow parameters that change slowly while a system runs: the internal resistance of a battery, the mass of a drone that picks up a load, the efficiency of each motor. The model is \(y_k=\varphi_k^{\mathsf T}\theta+e_k\), where \(\theta\) holds the unknown parameters and the **regressor** \(\varphi_k\) the known signals that multiply them. A battery's terminal voltage, \(V=V_{oc}-R_0I\), fits with \(\varphi=[1,\;-I]^{\mathsf T}\) and \(\theta=[V_{oc},\;R_0]^{\mathsf T}\).
Recursive least squares (RLS) is an online estimator of the parameters of a model that is linear in them, updating the least squares solution with each new sample without storing the history, and it is used to follow parameters that change slowly while a system runs: the internal resistance of a battery, the mass of a drone that picks up a load, the efficiency of each motor. The model is yk=φkTθ+eky_k=\varphi_k^{\mathsf T}\theta+e_kyk=φkTθ+ek, where θ\thetaθ holds the unknown parameters and the regressor φk\varphi_kφk the known signals that multiply them. A battery's terminal voltage, V=Voc−R0IV=V_{oc}-R_0IV=Voc−R0I, fits with φ=[1, −I]T\varphi=[1,;-I]^{\mathsf T}φ=[1,−I]T and θ=[Voc, R0]T\theta=[V_{oc},;R_0]^{\mathsf T}θ=[Voc,R0]T.
Kk=Pk−1φkλ+φkTPk−1φk,θ^k=θ^k−1+Kk(yk−φkTθ^k−1),Pk=1λ(Pk−1−KkφkTPk−1).\begin{gathered} K_k=\frac{P_{k-1}\varphi_k}{\lambda+\varphi_k^{\mathsf T}P_{k-1}\varphi_k},\qquad \hat\theta_k=\hat\theta_{k-1}+K_k\big(y_k-\varphi_k^{\mathsf T}\hat\theta_{k-1}\big),\\ P_k=\tfrac1\lambda\big(P_{k-1}-K_k\varphi_k^{\mathsf T}P_{k-1}\big). \end{gathered}Kk=λ+φkTPk−1φkPk−1φk,θ^k=θ^k−1+Kk(yk−φkTθ^k−1),Pk=λ1(Pk−1−KkφkTPk−1).
The middle line is the familiar correction: the old estimate plus a gain times the error in predicting the new sample. PPP is the parameters' covariance (up to a scale factor) and λ\lambdaλ, the forgetting factor, between 0.95 and 0.999, discounts old samples so the estimate can follow change; the effective memory is about 1/(1−λ)1/(1-\lambda)1/(1−λ) samples, a second of data for λ=0.99\lambda=0.99λ=0.99 at 100 Hz. RLS is a Kalman filter whose state is the parameter vector, with A=IA=IA=I and C=φkTC=\varphi_k^{\mathsf T}C=φkT, and forgetting plays the role of the process noise.
It only learns what the data excite. At constant current VocV_{oc}Voc and R0R_0R0 are indistinguishable: every sample says the same about their combination, and the current must vary for the two to separate. That is observability again, which identification calls excitation (persistent excitation).
Forgetting without excitation is a trap, covariance windup. While the system sits still, PPP grows like λ−k\lambda^{-k}λ−k; when a manoeuvre finally comes the gain is huge and the parameters jump. Bounding the trace of PPP or forgetting only in the directions the data excite are the usual guards.
When the parameters never change, one batch fit with all the data is simpler and better. RLS pays off when the estimate is needed while running, with bounded memory, and feeds a controller that redesigns itself from it (self-tuning regulator); echo cancellers in telephony run the same recursion on hundreds of filter taps.