control//adaptive control//dual control

Dual control is the formulation of adaptive control in which the controller explicitly splits its effort between regulating the plant and probing it to learn its unknown parameters, and it names the fundamental conflict of every system that learns while it acts. A good controller keeps the plant still; a still plant teaches nothing. Feldbaum formalized the problem around 1960: the ideal adaptive controller chooses each input both for its immediate effect on the error and for the information it will reveal, which pays off later in better control.


Dual control is the formulation of adaptive control in which the controller explicitly splits its effort between regulating the plant and probing it to learn its unknown parameters, and it names the fundamental conflict of every system that learns while it acts. A good controller keeps the plant still; a still plant teaches nothing. Feldbaum formalized the problem around 1960: the ideal adaptive controller chooses each input both for its immediate effect on the error and for the information it will reveal, which pays off later in better control.

A drone hovering with a constant command makes the point concrete. Its data contain one number, the ratio of thrust to weight; the mass and the thrust coefficient cannot be separated, and a friction term that appears only in motion is invisible (persistent excitation). A dual controller would sometimes move the vehicle slightly on purpose, accepting a small error now to know the plant better for the next manoeuvre, and would stop probing once the parameters were well known.

The exact solution is a stochastic dynamic program over the belief about the parameters (dynamic programming), intractable outside toy problems; it is the same structure as a partially observable decision problem (POMDP), where some actions are worth taking only because they clarify the state. It is the control version of the exploration-exploitation trade-off of reinforcement learning, and the opposite of certainty equivalence, which acts as if the current estimate were the truth.

Practice resolves it with cheap approximations: inject small excitation signals such as PRBS during commissioning or maintenance windows (experiment design), take advantage of natural manoeuvres (take-off, turns, setpoint changes), and freeze adaptation when the data carry no information (robust adaptation).

A fleet changes the economics. A new unit can start from the fleet's prior and needs far less probing to converge, because its siblings already explored (fleet).

Probing has a price in the field: on a process plant every excitation is off-spec product, and on an aircraft it is passenger comfort or margin, which is why probing is scheduled, small and bounded rather than continuous (adaptive control).