industrial//industrial data layer//CMMS

A CMMS (computerized maintenance management system) is the business software in which a plant records its assets, the maintenance planned for them and every work order carried out on them, and for anyone building a failure model it is where the labels live. Each order says which asset, what was done, which parts were used and when; IBM Maximo and the maintenance module of SAP are the usual large examples, under the wider name of enterprise asset management (EAM). It sits beside the historian, which holds the process signals, and the MES, which knows what was being produced, and shares neither their clock nor their identifiers.


A CMMS (computerized maintenance management system) is the business software in which a plant records its assets, the maintenance planned for them and every work order carried out on them, and for anyone building a failure model it is where the labels live. Each order says which asset, what was done, which parts were used and when; IBM Maximo and the maintenance module of SAP are the usual large examples, under the wider name of enterprise asset management (EAM). It sits beside the historian, which holds the process signals, and the MES, which knows what was being produced, and shares neither their clock nor their identifiers.

What a maintenance data scientist wants from it is a column saying failed here. What it holds is what a technician wrote, often in free text, with the date of the repair and not the date the fault began, and only for what somebody inspected. A pump that was opened and found worn has a record; the identical pump next to it that nobody opened has none, and its silence means not looked at, which is different from healthy.

The labels are biased in a known direction, and the direction matters more than the noise. Repairs are dated late, small faults fixed during another job go unrecorded, and machines replaced on a schedule before they failed never show their natural life (industrial labels, censored data).

Failures are rare in the record. With one failure per two thousand time windows, a model that always answers normal is right 99.95% of the time and useless, the trap of the base rate fallacy seen from the data side.

Joining it to the signals is the expensive part. The asset names of the CMMS rarely match the tags of the historian, and its dates are in local time, so linking an order to the vibration that preceded it is data contextualization work before it is modelling (predictive maintenance, preventive maintenance).