ML//MLOps//staged rollout
A staged rollout is a release procedure that delivers a new version of software, a model or a controller to a small fraction of the fleet first and widens it step by step, with health metrics that stop the release automatically and a rollback ready at every step, and it is how fleets of drones, cars, robots and plant controllers are updated without betting all of them at once. Typical steps are 1 % of the units, then 10 %, then all.
A staged rollout is a release procedure that delivers a new version of software, a model or a controller to a small fraction of the fleet first and widens it step by step, with health metrics that stop the release automatically and a rollback ready at every step, and it is how fleets of drones, cars, robots and plant controllers are updated without betting all of them at once. Typical steps are 1 % of the units, then 10 %, then all.
The reason is a failure that only a fleet can have. Hundreds of identical units running identical software fail together when the software is wrong: one bad update deployed everywhere is a common-mode failure, the case redundancy cannot protect against. A first step on a handful of units turns that into a local incident. The units in the first wave are best chosen to cover different operating conditions (climates, payloads, firmware revisions), since a fault that appears only in cold weather will not show on ten drones flying in July.
The stop rule is written before the release. It names the metrics (crash or failsafe rate, residuals of the estimator, CPU load, alarm counts), their thresholds and the observation time per stage, so that the rollout halts without a meeting. A metric chosen after the fact tends to be the one that looks fine.
Rollback must be cheap and tested. On embedded devices it rests on the OTA update scheme with two partitions: the unit boots the new image, runs its health checks, and returns to the old one by itself if they fail. A rollback path that has never been exercised is a second release waiting to go wrong.
It comes after shadow mode. Shadow mode answers whether the candidate decides well on real data without acting; the staged rollout answers whether it behaves well when it does act, which matters for controllers whose decisions change the states they will see next.
Software teams call the first small wave a canary, after the bird miners carried underground; the practice and the arithmetic are the same for a web service as for a fleet of pumps, with the difference that a bad drone update cannot be rolled back from a unit that has already crashed.