ML//MLOps//shadow mode

Shadow mode is a deployment stage in which a new model or controller runs in parallel with the one in charge, receives the same live inputs and logs what it would have decided, without its output reaching any actuator, and it is used to test a candidate on real field data before it is given control. The live system keeps flying the drone or running the line; the shadow only writes down its opinion.


Shadow mode is a deployment stage in which a new model or controller runs in parallel with the one in charge, receives the same live inputs and logs what it would have decided, without its output reaching any actuator, and it is used to test a candidate on real field data before it is given control. The live system keeps flying the drone or running the line; the shadow only writes down its opinion.

What makes it valuable is the comparison it allows. Each logged decision can be set against the decision the current system took and, later, against what actually happened: the bearing that failed, the anomaly that was real, the trajectory the drone flew. A new anomaly detector that would have raised forty alarms last week, thirty-eight of them on a healthy pump, is caught without a single false shutdown. A retrained fault classifier for a fleet of machines can run for a month across every unit, on every operating regime the fleet meets, which no test set gathered in advance covers.

Shadow mode measures the candidate on the data it will really face, at zero risk to the plant, but it never shows the candidate's own consequences. Its outputs are not applied, so the world it sees is the world the old system shaped.

That last limit matters most for controllers. A shadow controller sees states produced by the current controller; if it would have flown differently, the states that would have followed are never observed (the same open-loop limitation as log replay). Shadow mode therefore validates estimators, detectors and classifiers well and controllers only partly, which is why the next stage, a staged rollout on a small fraction of units with rollback ready, exists.

It costs compute and plumbing: the candidate must run on the deployed hardware with the deployed preprocessing (otherwise the test hides training-serving skew), and both decisions must be logged with timestamps so they can be compared with the outcome.

It fits the fleet case naturally: data gathered across many drones or machines are reidentified or retrained centrally, validated, run in shadow, then rolled out, the deployment loop of MLOps applied to control.