ML//MLOps

MLOps is the software engineering of a learned model's whole life (versioning, deployment, monitoring and retraining), and it is what turns a model that works on a laptop into one that works on 300 machines for five years. A trained model is code plus weights plus the data that shaped them, so its behaviour can change without anyone touching a line of code: a sensor supplier changes, the summer arrives, a preprocessing step is rewritten. MLOps exists because of that third dependency.


MLOps is the software engineering of a learned model's whole life (versioning, deployment, monitoring and retraining), and it is what turns a model that works on a laptop into one that works on 300 machines for five years. A trained model is code plus weights plus the data that shaped them, so its behaviour can change without anyone touching a line of code: a sensor supplier changes, the summer arrives, a preprocessing step is rewritten. MLOps exists because of that third dependency.

The famous drawing of hidden technical debt in ML (Sculley and colleagues, 2015) shows the model code as a small box surrounded by huge boxes of data collection, feature extraction, configuration, serving and monitoring. In an industrial setting the expensive part is the feature pipeline: computing the same features, the same way and on time, on the machine that runs the model as on the one that trained it. A gradient-boosted classifier that takes microseconds to evaluate can need months of plumbing around it.

Version everything that changes behaviour: code, weights, a snapshot of the training data, configuration, the calibration of each unit and the firmware of its sensors. The acid test is a question with a unit number in it: can you rebuild exactly the model running on unit 217?

Getting a model into the field safely is a ladder of rehearsals. In shadow mode the new model computes beside the live one and never acts; a staged rollout then hands it control on 1 %, 10 % and finally all units, with metrics that stop the rollout by themselves; on embedded devices the mechanism underneath is an OTA update with two partitions and automatic rollback.

The model that was validated is often not the model that runs. Training-serving skew covers the quiet ways they differ (preprocessing, an int8 conversion, a prototype sensor), and the remedy is to validate the deployed artifact itself.

Once deployed, the world moves away from the training set. Data drift is watched on the inputs and on residuals, because in predictive maintenance the labels that would tell you the model is wrong arrive weeks later, when a machine fails or does not.

A model in production is more like a pet than a piece of furniture: it needs feeding (data), the vet (retraining and revalidation) and a caretaker who stays after the person who brought it has left. That recurring cost belongs in its total cost of ownership from the first day, and it is why a threshold maintained by adjusting one number often wins; for models served from the cloud, the same telemetry that watches quality also watches GPU utilization and cost per request (FinOps).

The same discipline applies to controllers and estimators as much as to networks: an adaptive gain table or a retuned Kalman filter pushed to a fleet carries the same risks, the first of them a common-mode failure when one bad update reaches every unit at once.