Survival › Time-to-event

Kaplan-Meier estimator

The Kaplan-Meier estimator reconstructs the survival function from right-censored follow-up data, with an optional log-rank test for differences between groups.

What is Kaplan-Meier estimator?

Time-to-event data has two complications standard regression can't handle: censoring (some subjects haven't had the event by end-of-followup) and skewed time distributions. The Kaplan-Meier estimator handles both via a step function that drops at each observed event by a factor (1 − dᵢ/nᵢ) where dᵢ is events and nᵢ is at-risk subjects at time tᵢ.

The log-rank test compares survival between groups by aggregating the per-event hypergeometric residuals. It has maximum power when the hazards are proportional across time; when they cross, log-rank can be insensitive and RMST or restricted mean differences are better.

K-M is purely descriptive — it summarises what happened, not why. For covariate-adjusted comparisons use Cox PH. For competing-risks data (where multiple types of events can preclude the event of interest), K-M overestimates cumulative incidence and you should use the cumulative-incidence function (Fine-Gray-style).

When should I use Kaplan-Meier estimator?

  • Visualising survival / time-to-event distributions by group.
  • Standard non-parametric first look before fitting Cox PH or a parametric model.
  • Median survival reporting with CIs (when the curve drops below 0.5).

What data does it need?

Time + event (0/1) + optional grouping factor.

What does it report?

Step survival curve with CI shading, median survival per group, log-rank p, pairwise comparisons with Holm-adjusted p.

What does it assume?

  • Independent observations.
  • Censoring is non-informative (subjects who censor aren't at systematically different risk).
  • For the log-rank test only: proportional hazards across groups.

Formula

Ŝ(t) = Π_{tᵢ ≤ t} (1 − dᵢ / nᵢ)

How do I interpret the result?

Median survival is the time at which Ŝ(t) crosses 0.5. If the curve never drops that far, median isn't estimable from the data — report a different landmark (e.g. 5-year survival) instead.

Crossing K-M curves between groups ⇒ non-proportional hazards ⇒ log-rank is underpowered. RMST or landmark analysis is the better summary in that case.

See also

References

  • Kaplan & Meier (1958). Nonparametric estimation from incomplete observations. JASA 53(282).