mathematics//statistics//survival analysis//Kaplan-Meier estimator
The Kaplan-Meier estimator is a nonparametric estimate of the survival function \(S(t)\), the fraction of a population still free of an event after time \(t\), built directly from durations of which many are censored, and it is the first plot drawn of any life dataset: bearings in a fleet of pumps, batteries in a fleet of drones, patients in a trial. It assumes no shape for the curve, which makes it the honest reference against which a Weibull distribution or any other model is later judged (survival analysis).
The Kaplan-Meier estimator is a nonparametric estimate of the survival function S(t)S(t)S(t), the fraction of a population still free of an event after time ttt, built directly from durations of which many are censored, and it is the first plot drawn of any life dataset: bearings in a fleet of pumps, batteries in a fleet of drones, patients in a trial. It assumes no shape for the curve, which makes it the honest reference against which a Weibull distribution or any other model is later judged (survival analysis).
The idea is to chain conditional survivals. At each time a failure occurs, the share of the units still under observation (the risk set) that survived that instant is computed, and the curve is the running product:
S^(t)=∏ti≤t(1−dini),\hat S(t)=\prod_{t_i\le t}\left(1-\frac{d_i}{n_i}\right),S^(t)=ti≤t∏(1−nidi),
with did_idi the failures at time tit_iti and nin_ini the units at risk just before it. A censored unit (replaced while healthy, or still running) counts in every risk set up to its last known age and then leaves without pulling the curve down. Ten bearings: one fails at 400 h, so S^=9/10=0.9\hat S=9/10=0.9S^=9/10=0.9; one is replaced healthy at 500 h and leaves; at 700 h another fails among the eight still at risk, so S^=0.9×7/8≈0.79\hat S=0.9\times7/8\approx0.79S^=0.9×7/8≈0.79. Counting the replaced bearing as a failure would have given 0.8 at 500 h and 0.7 at 700 h, a pessimism that grows with every preventive replacement.
The curve is a staircase that drops only at failures and gets less certain to the right.
Each step is estimated from the units still at risk, and at long ages those are few, so the right end of the curve rests on a handful of units; its confidence band (Greenwood's formula) widens there, and a flat tail may only mean nobody was left to fail.
It is valid when censoring carries no information about the unit's health. If bearings are pulled early because they sounded bad, the censored units were the weak ones, the survivors look healthier than the population, and the curve is optimistic; the maintenance policy that produced the data has to be known before the curve is read.
It compares groups well and explains them poorly. Two curves (supplier A against supplier B, hot site against cool site) can be drawn and tested for a difference (the log-rank test); how much each factor shortens life, with several factors at once, is the job of the Cox model or a parametric fit by maximum likelihood.
In Python it is one call, KaplanMeierFitter in the lifelines library, given the durations and a flag per unit saying whether it failed or was censored.