mathematics//decision theory//certainty equivalence

Certainty equivalence is the design principle of estimating the unknown state or parameters of a system and then deciding, or computing the controller, as if that estimate were exactly true; it is how nearly every real autopilot, adaptive controller and robot planner bridges estimation and action. The estimator runs on one side, the controller or planner on the other, and the only thing passed between them is the best guess. (Economics uses *certainty equivalent* for something else, the sure amount a person values as much as a gamble.)


Certainty equivalence is the design principle of estimating the unknown state or parameters of a system and then deciding, or computing the controller, as if that estimate were exactly true; it is how nearly every real autopilot, adaptive controller and robot planner bridges estimation and action. The estimator runs on one side, the controller or planner on the other, and the only thing passed between them is the best guess. (Economics uses certainty equivalent for something else, the sure amount a person values as much as a gamble.)

It is the shortcut industry takes instead of solving a POMDP, where decisions would be made on the whole belief. For one important case it is no shortcut at all: with linear dynamics, Gaussian noise and quadratic cost (LQG), feeding the Kalman filter's estimate to the LQR gain designed for perfect state knowledge is exactly optimal, which is the separation principle. In adaptive control it is the logic of the self-tuning regulator of Åström and Wittenmark (1973): identify the model's parameters online with recursive least squares and recompute the controller at every step as if they were right. PX4's hover-thrust estimator is the humblest example in production, a one-parameter estimate of thrust efficiency that rescales the feedforward as if it were exact.

Certainty equivalence works while the uncertainty is small against the consequences.

A robot that does not know whether it is 30 cm or 3 cm from a step should not act as if it stood at the mean; with safety margins sized from the estimate's covariance, it is industry practice.

It never acts to learn. A controller that treats its estimate as truth sees no value in probing, so it can sit on a wrong estimate whenever its own actions keep the system too still to correct it; dual control adds the probing that certainty equivalence leaves out (exploration-exploitation trade-off), and active perception does the same for a planner.

A bad start feeds itself. With a poor initial estimate the controller is poor, and the data it generates are poorer still, so adaptive schemes start from good prior values and bound the parameters they let the estimator reach.