ML//applied ML//industrial labels

Industrial labels are the answers a supervised model learns from in a plant (this pump failed, this part was scrap, this alarm was real), and they are rarely measured facts: a failure label is what a technician wrote in a work order, dated at the repair rather than at the onset, and only for what was inspected. Every failure-prediction project depends on them, and most of those that stall, stall here.


Industrial labels are the answers a supervised model learns from in a plant (this pump failed, this part was scrap, this alarm was real), and they are rarely measured facts: a failure label is what a technician wrote in a work order, dated at the repair rather than at the onset, and only for what was inspected. Every failure-prediction project depends on them, and most of those that stall, stall here.

They come from the maintenance management system (CMMS) as free text: a cause field that says fixed, two failure types logged under one code, half the failure dates out of step with the sensor data. The sensor history lives elsewhere, in the historian (often compressed with a deadband, so what lies between samples is an interpolation nobody measured), in the event logs of the PLC and the SCADA, and in the MES that knows which product and batch was running. Building the training table means aligning all of it in time across clocks that disagree and daylight-saving changes (time synchronization), and that phase eats most of the project.

Industrial labels are late, censored and imbalanced.

Late, because the repair date trails the onset by days or weeks. Censored, because a machine that has not failed yet is not known to be healthy, only to have survived so far (censored data). Imbalanced, because with one failure per 2,000 windows a model that always answers normal is 99.95 % accurate (base rate, class imbalance). A result that looks too good is a leak until proven otherwise.

Label noise is the polite name for all of this, and its errors follow patterns. Failures of inspected components are over-represented, the onset is shifted late, and the most dramatic failures get the most careful records. A model trained on such labels learns the habits of the maintenance team along with the physics of the machine.

Events count more than rows. With a failure rate of 0.1 % and 10,000 records there are about ten failures, little to train on and almost nothing to validate with (sample size).

Work orders are the bottleneck. What separates the projects that reach the plant from the demos is rarely the algorithm; it is the maintenance records. Fixing the form (a coded cause, the date the symptom was first seen) often buys more than any model, and a rule written with the plant expert may solve most of the problem before anyone labels 10,000 examples (problem framing, predictive maintenance).