Survival analysis and censored outcomes
Survival analysis handles the patients you had to stop watching before you saw the outcome, instead of quietly pretending they never existed.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Survival analysis answers "how long until this happens", when some people leave the study before it happens to them.
Think of a race where you must publish results at a fixed time. Not every runner has crossed the finish line yet. Some runners are still out on the course, visibly still running. You cannot honestly call them "losers" — you do not yet know their final time.
Medical studies hit this constantly. A patient starts treatment. The study ends, or the patient moves away, before anyone learns whether they relapsed. That patient is not "cured." They are censored — meaning their true outcome is unknown, not that nothing happened.
Why it exists
A tempting shortcut is to average how long patients went without relapsing, and call that the answer. That shortcut quietly assumes every censored patient relapsed the moment the study stopped watching. That is almost certainly false, and it biases the answer toward "sicker than reality."
The opposite shortcut — dropping every censored patient from the analysis — is equally broken. It throws away the patients who were doing best. Those who survived longest are also most likely to still be alive and unrelapsed when the study ends.
Survival analysis was built to use every patient's information honestly. That means the ones who relapsed, for as long as they lasted. It also means the ones who did not, for as long as they were watched.
How it works
Patient | Followed for | What happened
---------|------------------|------------------
A | 8 months | relapsed (event observed)
B | 8 months | still fine when study ended (censored)
C | 24 months | relapsed (event observed)
D | 24 months | still fine when study ended (censored)A survival curve starts at 100% — "everyone is relapse-free" — and steps downward only when a real relapse happens. Censored patients do not pull the curve down — they stop contributing information from the moment they leave.
Where you have already seen it
- A phone's battery life estimate ("about 6 more hours") is a survival-style prediction. Some usage patterns drain it sooner, some later.
- A warranty period on a product is set using survival analysis on how long similar products lasted before failing.
An honest warning
Survival analysis assumes that why a patient was censored has nothing to do with their real outcome. If patients tend to drop out of a study because they are getting sicker, that assumption breaks. The resulting curve then looks too optimistic. This is a real and common failure in medical studies. It is worth asking about, every time you see a survival curve.
Remember this
- A censored patient's true outcome is unknown, not "no event happened."
- Averaging raw follow-up times, or dropping censored patients, both bias the result.
- A survival curve only drops at times a real event was observed — censoring alone never moves it.
What to learn next
- Confidence intervals — the general uncertainty concept survival curves build on.
- What is time series data? — a related but different way of thinking about data over time.
- Validating a model before it touches patients — where a model's predicted survival curve has to earn trust before deployment.
Developer — Code and libraries.
Setup
pip install lifelines pandaslifelines is the standard Python library for survival analysis — actively maintained, and used across both academic and industry work in this area.
Minimal runnable code
import pandas as pd
from lifelines import KaplanMeierFitter
# One row per patient after starting a treatment. "duration" is months
# followed. "relapsed" is 1 if we saw the event, 0 if the patient was still
# relapse-free when we last saw them (they left the study, or it ended).
data = pd.DataFrame({
"duration": [5, 8, 8, 12, 12, 15, 18, 20, 22, 24, 24, 30],
"relapsed": [1, 1, 0, 1, 0, 1, 0, 1, 1, 0, 0, 1],
})
# The naive move: average every duration and call it "typical time to relapse".
naive_average = data["duration"].mean()
print(f"naive average of all durations: {naive_average:.1f} months")
print("that treats every censored patient as if they relapsed the day we stopped watching.")
print()
km = KaplanMeierFitter()
km.fit(durations=data["duration"], event_observed=data["relapsed"])
print(f"Kaplan-Meier median relapse-free survival: {km.median_survival_time_} months")naive average of all durations: 16.5 months that treats every censored patient as if they relapsed the day we stopped watching. Kaplan-Meier median relapse-free survival: 22.0 months
What actually happened
The naive average, 16.5 months, and the Kaplan-Meier median, 22.0 months, disagree by more than five months — over a third of the naive number. That gap is not noise. It is the cost of treating censored patients as if they had already failed.
The Kaplan-Meier estimator only lets the curve drop at the exact months a relapse was actually observed: 5, 8, 12, 15, 20, 22, 30. Censored patients — at months 8, 12, 24, and 24 — quietly leave the "at risk" group without dragging the curve down, because leaving does not mean relapsing.
Line by line, the parts that are not obvious:
event_observedtakes 1 for a real event, 0 for censoring — get this backwards and every conclusion in the analysis inverts silently, with no error raised.km.median_survival_time_is the duration at which the survival curve first crosses 50% — not the average of the raw durations, and not always defined if the curve never crosses 50% within the observed data.KaplanMeierFitterhandles right-censoring specifically — patients who leave after some follow-up, which is the overwhelming majority of censoring in real medical studies.
Common mistakes
Treating a 0 in event_observed as "the patient is fine". It means "we do not know" — the patient might relapse the following week, after the data was collected.
Comparing two groups' raw average durations instead of their survival curves. This is a routine mistake even in published work, and it is exactly the bias shown above.
Forgetting that censoring must be non-informative. If sicker patients are more likely to drop out of a study, the survival curve looks better than reality. Kaplan-Meier has no way to detect this on its own — it is an assumption the study design has to earn.
Try it yourself
Change patient at index 5 (duration 15, relapsed) to duration=15, relapsed=0 — imagine they were actually censored, not observed relapsing. Refit and watch the median survival time shift, even though only one row changed.
What to learn next
- Confidence intervals — the general uncertainty machinery a survival curve's confidence band relies on.
- When the radiologists disagree — a different kind of "we don't actually know the true label" problem.
- Predicting patient deterioration — a related but distinct clinical prediction task, on a shorter timescale.
Researcher — Mathematics and papers.
The Kaplan-Meier estimator
For observed distinct event times t_1 < t_2 < ... < t_k, the Kaplan-Meier estimate of the survival function is:
S_hat(t) = PRODUCT over all t_i <= t of ( 1 - d_i / n_i )S_hat(t)— estimated probability of surviving (remaining event-free) past timetd_i— number of events observed exactly at timet_in_i— number of subjects still "at risk" (neither had the event nor were censored) immediately beforet_i
Each factor (1 - d_i / n_i) is the estimated probability of surviving that specific event time, given survival up to it. The product accumulates these conditional probabilities — this is the nonparametric maximum likelihood estimator of the survival function under right-censoring (Kaplan & Meier, 1958).
The hazard function
The hazard function h(t) is the instantaneous event rate at time t, conditional on survival to t:
h(t) = lim_{dt -> 0} P(t <= T < t + dt | T >= t) / dtT— the random variable representing time-to-event
Survival and hazard are related by S(t) = exp( -integral from 0 to t of h(u) du ). The Cox proportional hazards model (Cox, 1972) extends this to incorporate covariates x, assuming each covariate multiplies the baseline hazard by a constant factor:
h(t | x) = h_0(t) * exp( beta^T x )h_0(t)— an unspecified baseline hazard, common to all subjectsbeta— a coefficient vector, estimated via partial likelihood without ever needing to specifyh_0(t)directly
This semi-parametric property — no assumption about the shape of h_0(t) — is why Cox regression remains the dominant model for covariate-adjusted survival analysis in medicine, in preference to fully parametric alternatives.
Evaluation: the concordance index
Model discrimination for survival predictions is measured with the concordance index (C-index): among all comparable pairs of subjects, the fraction where the subject predicted to have higher risk actually experienced the event first. It generalises ROC-AUC to the presence of censoring, and reduces exactly to ROC-AUC in the special case of no censoring at all.
Cost
Kaplan-Meier estimation is O(n log n), dominated by sorting event times. Cox regression via Newton-Raphson on the partial likelihood is O(n^2 d) per iteration in the naive implementation, though production libraries use algorithmic refinements that scale considerably better in practice for large n.
Key references
- Kaplan, E. & Meier, P. (1958). Nonparametric Estimation from Incomplete Observations. JASA 53(282). The original estimator.
- Cox, D. (1972). Regression Models and Life-Tables. JRSS B 34(2). Proportional hazards and partial likelihood.
- Harrell, F. et al. (1982). Evaluating the Yield of Medical Tests. JAMA. Introduces the concordance index.
- Katzman, J. et al. (2018). DeepSurv: Personalized Treatment Recommender System Using a Cox Proportional Hazards Deep Neural Network. BMC Medical Research Methodology. A neural extension of Cox regression.
Current state
Classical Cox regression remains the default in clinical research due to interpretability and regulatory familiarity. Neural survival models (DeepSurv, and later transformer-based variants) can better capture non-linear and non-proportional-hazards relationships, at the cost of the interpretability that makes Cox models easy to explain to a clinical review board — a trade-off directly relevant to validating a model before it touches patients.
What to learn next
- Confidence intervals — the statistical foundation survival curve uncertainty bands build on.
- Model evaluation — how evaluation metrics generalise (and do not generalise) to censored outcomes.
- Validating a model before it touches patients — what comes after a survival model looks good on paper.