AI in Healthcare

Survival analysis and censored outcomes

Survival analysis handles the patients you had to stop watching before you saw the outcome, instead of quietly pretending they never existed.

Read these first

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. An honest warning
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Survival analysis answers "how long until this happens", when some people leave the study before it happens to them.

Think of a race where you must publish results at a fixed time. Not every runner has crossed the finish line yet. Some runners are still out on the course, visibly still running. You cannot honestly call them "losers" — you do not yet know their final time.

Medical studies hit this constantly. A patient starts treatment. The study ends, or the patient moves away, before anyone learns whether they relapsed. That patient is not "cured." They are censored — meaning their true outcome is unknown, not that nothing happened.

Why it exists

A tempting shortcut is to average how long patients went without relapsing, and call that the answer. That shortcut quietly assumes every censored patient relapsed the moment the study stopped watching. That is almost certainly false, and it biases the answer toward "sicker than reality."

The opposite shortcut — dropping every censored patient from the analysis — is equally broken. It throws away the patients who were doing best. Those who survived longest are also most likely to still be alive and unrelapsed when the study ends.

Survival analysis was built to use every patient's information honestly. That means the ones who relapsed, for as long as they lasted. It also means the ones who did not, for as long as they were watched.

How it works

Patient  |  Followed for   |  What happened
---------|------------------|------------------
   A     |     8 months     |  relapsed  (event observed)
   B     |     8 months     |  still fine when study ended  (censored)
   C     |    24 months     |  relapsed  (event observed)
   D     |    24 months     |  still fine when study ended  (censored)

A survival curve starts at 100% — "everyone is relapse-free" — and steps downward only when a real relapse happens. Censored patients do not pull the curve down — they stop contributing information from the moment they leave.

Where you have already seen it

  • A phone's battery life estimate ("about 6 more hours") is a survival-style prediction. Some usage patterns drain it sooner, some later.
  • A warranty period on a product is set using survival analysis on how long similar products lasted before failing.

An honest warning

Survival analysis assumes that why a patient was censored has nothing to do with their real outcome. If patients tend to drop out of a study because they are getting sicker, that assumption breaks. The resulting curve then looks too optimistic. This is a real and common failure in medical studies. It is worth asking about, every time you see a survival curve.

Remember this

  • A censored patient's true outcome is unknown, not "no event happened."
  • Averaging raw follow-up times, or dropping censored patients, both bias the result.
  • A survival curve only drops at times a real event was observed — censoring alone never moves it.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install lifelines pandas

lifelines is the standard Python library for survival analysis — actively maintained, and used across both academic and industry work in this area.

Minimal runnable code

survival_demo.py
import pandas as pd
from lifelines import KaplanMeierFitter

# One row per patient after starting a treatment. "duration" is months
# followed. "relapsed" is 1 if we saw the event, 0 if the patient was still
# relapse-free when we last saw them (they left the study, or it ended).
data = pd.DataFrame({
    "duration": [5, 8, 8, 12, 12, 15, 18, 20, 22, 24, 24, 30],
    "relapsed": [1, 1, 0, 1, 0, 1, 0, 1, 1, 0, 0, 1],
})

# The naive move: average every duration and call it "typical time to relapse".
naive_average = data["duration"].mean()
print(f"naive average of all durations: {naive_average:.1f} months")
print("that treats every censored patient as if they relapsed the day we stopped watching.")
print()

km = KaplanMeierFitter()
km.fit(durations=data["duration"], event_observed=data["relapsed"])
print(f"Kaplan-Meier median relapse-free survival: {km.median_survival_time_} months")
Output
naive average of all durations: 16.5 months
that treats every censored patient as if they relapsed the day we stopped watching.

Kaplan-Meier median relapse-free survival: 22.0 months

What actually happened

The naive average, 16.5 months, and the Kaplan-Meier median, 22.0 months, disagree by more than five months — over a third of the naive number. That gap is not noise. It is the cost of treating censored patients as if they had already failed.

The Kaplan-Meier estimator only lets the curve drop at the exact months a relapse was actually observed: 5, 8, 12, 15, 20, 22, 30. Censored patients — at months 8, 12, 24, and 24 — quietly leave the "at risk" group without dragging the curve down, because leaving does not mean relapsing.

Line by line, the parts that are not obvious:

  • event_observed takes 1 for a real event, 0 for censoring — get this backwards and every conclusion in the analysis inverts silently, with no error raised.
  • km.median_survival_time_ is the duration at which the survival curve first crosses 50% — not the average of the raw durations, and not always defined if the curve never crosses 50% within the observed data.
  • KaplanMeierFitter handles right-censoring specifically — patients who leave after some follow-up, which is the overwhelming majority of censoring in real medical studies.

Common mistakes

Treating a 0 in event_observed as "the patient is fine". It means "we do not know" — the patient might relapse the following week, after the data was collected.

Comparing two groups' raw average durations instead of their survival curves. This is a routine mistake even in published work, and it is exactly the bias shown above.

Forgetting that censoring must be non-informative. If sicker patients are more likely to drop out of a study, the survival curve looks better than reality. Kaplan-Meier has no way to detect this on its own — it is an assumption the study design has to earn.

Try it yourself

Change patient at index 5 (duration 15, relapsed) to duration=15, relapsed=0 — imagine they were actually censored, not observed relapsing. Refit and watch the median survival time shift, even though only one row changed.

What to learn next

Researcher — Mathematics and papers.

The Kaplan-Meier estimator

For observed distinct event times t_1 < t_2 < ... < t_k, the Kaplan-Meier estimate of the survival function is:

S_hat(t) = PRODUCT over all  t_i <= t   of   ( 1 - d_i / n_i )
  • S_hat(t) — estimated probability of surviving (remaining event-free) past time t
  • d_i — number of events observed exactly at time t_i
  • n_i — number of subjects still "at risk" (neither had the event nor were censored) immediately before t_i

Each factor (1 - d_i / n_i) is the estimated probability of surviving that specific event time, given survival up to it. The product accumulates these conditional probabilities — this is the nonparametric maximum likelihood estimator of the survival function under right-censoring (Kaplan & Meier, 1958).

The hazard function

The hazard function h(t) is the instantaneous event rate at time t, conditional on survival to t:

h(t) = lim_{dt -> 0}  P(t <= T < t + dt | T >= t) / dt
  • T — the random variable representing time-to-event

Survival and hazard are related by S(t) = exp( -integral from 0 to t of h(u) du ). The Cox proportional hazards model (Cox, 1972) extends this to incorporate covariates x, assuming each covariate multiplies the baseline hazard by a constant factor:

h(t | x) = h_0(t) * exp( beta^T x )
  • h_0(t) — an unspecified baseline hazard, common to all subjects
  • beta — a coefficient vector, estimated via partial likelihood without ever needing to specify h_0(t) directly

This semi-parametric property — no assumption about the shape of h_0(t) — is why Cox regression remains the dominant model for covariate-adjusted survival analysis in medicine, in preference to fully parametric alternatives.

Evaluation: the concordance index

Model discrimination for survival predictions is measured with the concordance index (C-index): among all comparable pairs of subjects, the fraction where the subject predicted to have higher risk actually experienced the event first. It generalises ROC-AUC to the presence of censoring, and reduces exactly to ROC-AUC in the special case of no censoring at all.

Cost

Kaplan-Meier estimation is O(n log n), dominated by sorting event times. Cox regression via Newton-Raphson on the partial likelihood is O(n^2 d) per iteration in the naive implementation, though production libraries use algorithmic refinements that scale considerably better in practice for large n.

Key references

  • Kaplan, E. & Meier, P. (1958). Nonparametric Estimation from Incomplete Observations. JASA 53(282). The original estimator.
  • Cox, D. (1972). Regression Models and Life-Tables. JRSS B 34(2). Proportional hazards and partial likelihood.
  • Harrell, F. et al. (1982). Evaluating the Yield of Medical Tests. JAMA. Introduces the concordance index.
  • Katzman, J. et al. (2018). DeepSurv: Personalized Treatment Recommender System Using a Cox Proportional Hazards Deep Neural Network. BMC Medical Research Methodology. A neural extension of Cox regression.

Current state

Classical Cox regression remains the default in clinical research due to interpretability and regulatory familiarity. Neural survival models (DeepSurv, and later transformer-based variants) can better capture non-linear and non-proportional-hazards relationships, at the cost of the interpretability that makes Cox models easy to explain to a clinical review board — a trade-off directly relevant to validating a model before it touches patients.

What to learn next