industrial//maintenance//prognostics//remaining useful life
The remaining useful life of a unit is the time it has left before it can no longer perform its function, \(T-t\) with \(T\) its failure time, conditioned on everything known about it at time \(t\), and it is the output of prognostics that maintenance decisions are built on. Because the future degradation, the load and even the definition of failure are uncertain, it is a probability distribution over time.
The remaining useful life of a unit is the time it has left before it can no longer perform its function, T−tT-tT−t with TTT its failure time, conditioned on everything known about it at time ttt, and it is the output of prognostics that maintenance decisions are built on. Because the future degradation, the load and even the definition of failure are uncertain, it is a probability distribution over time.
There are three ways to estimate it, from least to most knowledge of the unit. By age: the probability of lasting τ\tauτ longer having reached ttt is R(t+τ)/R(t)R(t+\tau)/R(t)R(t+τ)/R(t), with RRR the population's survival function from reliability, so every unit of the same age gets the same prognosis. By degradation: a health indicator, a degradation model of its evolution and a failure threshold zFz_FzF, projected forward by simulation. By direct regression: a model learns the map from sensor histories to remaining life from many units run to failure, the famous benchmark being NASA's C-MAPSS turbofan simulations, whose scoring function punishes late predictions more than early ones, as reality does.
median remaining life781 h 90 % interval522 to 1224 h failure before the stop1.0 % decisionwait for the stop After 250 h of readings, the remaining life has a median of 781 h and a 90 % interval from 522 h to 1224 h. The chance of failing before the stop planned in 450 h is 1.0 %, within the accepted risk of 5 %: wait for the stop.
With few hours of data the fan of projected trajectories and the histogram of failure times are wide; slide today forward and the 90 % interval shrinks. Push the planned stop later until the share of failures before it exceeds the acceptable risk, and the decision flips to intervene.
The decision depends on the left tail of the distribution.
A prognosis stated as three hundred hours left is dangerous: what matters is the probability of failing before the next opportunity to replace the part. A median of 500 hours with a stop in 300 can still carry a 20 % risk.
The width of the distribution has four sources: the noise in the current estimate of the indicator; the degradation model (linear and exponential fit early data equally and diverge late); the future load, the forgotten one, which is not in the past data; and the failure threshold, more convention (stop at 11 mm/s) than law. Near the end the distribution narrows; far from it, it is wide and skewed.
Direct regression needs what a well-maintained plant lacks, many complete histories to failure, and models validated by rows instead of by unit leak information between the histories of one engine; split by unit.
The probability of failing before the stop, read from this distribution, is the input of the cost rule in predictive maintenance; designing on the low quantile instead of the mean is one case of tails over means.