control//system identification
System identification is the discipline of building a dynamic model of a system (its matrices \(A\) and \(B\), a transfer function, or a nonlinear \(f\)) from measured input and output data, and it is the door through which data enters control: every LQR gain, every MPC prediction and every PID rule rests on a model that someone identified. It is machine learning with one extra condition: the result must be a dynamical system good enough to estimate and control with.
System identification is the discipline of building a dynamic model of a system (its matrices AAA and BBB, a transfer function, or a nonlinear fff) from measured input and output data, and it is the door through which data enters control: every LQR gain, every MPC prediction and every PID rule rests on a model that someone identified. It is machine learning with one extra condition: the result must be a dynamical system good enough to estimate and control with.
The work has four steps, and most failures happen in the first and the last. Design the experiment: which input signal, what amplitude, how long (experiment design). Choose the structure: white box (physics with measured parameters), grey box (physics with fitted parameters such as mass, inertia or thrust constant, the default in industry, grey-box model) or black box (ARX models, state-space models found by subspace methods through an SVD of data matrices, neural networks). Fit: least squares when the model is linear in its parameters, the general prediction error method otherwise. Validate on data not used for fitting, in free-run simulation, and check that the residuals are white (model validation).
Identifying is designing an experiment, choosing a structure, fitting and validating in free run.
Fitting is the easy step. A model fitted to data in which the system did not move tells you nothing about its dynamics, and a model validated one step ahead can look excellent while having learned only that the next sample resembles the last.
Identifiability is whether the parameters can be deduced from the data at all; it requires that the system moved enough while observed (persistent excitation). It is the third of three neighbours not to mix: measurable (a sensor gives it), observable (the state can be reconstructed, observability), identifiable (the parameters can be learned).
Feedback hides causation. In a well-regulated oven, heater power and temperature barely correlate, because the loop spends the power cancelling open doors and cold loads; a naive analysis concludes that the heater has no effect. To learn a cause, intervene with a step or an excitation signal.
Closed-loop identification is unavoidable for unstable plants such as a drone, which cannot fly open loop: the excitation is added to the setpoint. Feedback correlates the input with the noise and biases naive least squares, so methods that account for the loop are needed.
For most industrial PID loops a step test and three numbers suffice (FOPDT model); drone autopilots do the same in miniature with their autotune modes, which excite the attitude, fit a simple model and compute gains. Higher-order models or networks pay off for MPC or serious simulation.
Online, continuously, identification becomes indirect adaptive control (self-tuning regulator). Typical tools: SysIdentPy, python-control, scipy.optimize.least_squares for grey boxes.