mathematics//statistics//survival analysis//censored data

Censored data are observations of a duration whose end was not seen, only a bound on it: a unit still running when the data were collected, or replaced before it failed, is known to last at least its current age. They are the normal state of life data in a plant, where failures are few and most units are either still in service or were changed preventively, and using them correctly is the difference between a life estimate and a fiction. Reliability engineers call them suspensions.


Censored data are observations of a duration whose end was not seen, only a bound on it: a unit still running when the data were collected, or replaced before it failed, is known to last at least its current age. They are the normal state of life data in a plant, where failures are few and most units are either still in service or were changed preventively, and using them correctly is the difference between a life estimate and a fiction. Reliability engineers call them suspensions.

The common case is right censoring, the true time lies somewhere to the right of what was recorded. Its source is often the maintenance policy itself. A fleet that replaces every bearing at 8,000 hours never sees one fail at 12,000; every life in its records is either a failure before 8,000 hours or a bearing known to have reached 8,000, and nothing describes what happens after. Treating those removals as failures, or dropping them, misreads the policy as the physics.

Discarding survivors makes components look short-lived.

Five bearings failed between 410 and 1,130 hours, and fifteen are still running at 1,200 hours. Fitting a Weibull distribution to the five failures alone gives a characteristic life of about 860 hours; including the fifteen as censored observations gives about 2,100. The first answer would have every bearing replaced at mid-life, with full statistical confidence. This survivor exclusion bias is the textbook selection bias: only the units that died early were looked at.

The methods that handle censoring are those of survival analysis: the Kaplan-Meier curve without a model, maximum likelihood with censored terms for a parametric one. In SciPy, stats.CensoredData(uncensored=failures, right=survivors) passed to weibull_min.fit is all it takes.

Left censoring (the event happened before the first inspection, at an unknown time) and interval censoring (it happened between two inspections) appear with periodic checks; they need their own terms and are routinely, and wrongly, recorded as if the event occurred at the inspection.

Censoring breaks naive machine learning in the same way. A machine that has not failed yet is a life whose end is unknown, and labelling it as healthy teaches a model that failures are rarer and later than they are (industrial labels).