Causal Inference Basics

Colliders and selection bias

A collider is a variable caused by two others — and the moment you filter or adjust on it, you manufacture a relationship between causes that never touched each other.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A collider is a variable that two other things both push on. Look only at the cases where the collider "happened", and a false link appears between those two things.

A college admits students who shine at academics or at sports — either is enough. Now walk around campus and study the admitted students. The brilliant scholars are, oddly, rarely great athletes. The star athletes are, oddly, rarely toppers. On campus, books and sport look like enemies.

Nothing in the world connects them — in the whole population of applicants, they are unrelated. The admission gate did it. A student weak in one dimension is on campus because they were strong in the other. Filtering through the gate manufactured the pattern.

Why it exists

The DAG lesson had two variable types: confounders (adjust) and mediators (leave alone). The collider is the third: a variable with two arrows pointing into it. Admission has arrows in from academics and from sports.

Colliders are harmless while ignored. The damage begins when you condition on one. That means filtering your data through it, studying only one side of it, or adding it to a regression. Think of every dataset that has an entry gate: surviving customers, admitted patients, published papers, hired candidates. Each was filtered through a collider before you ever opened it.

How it works

academics  →  ADMISSION  ←  sports        (two arrows IN: a collider)

whole applicant pool:  academics and sports unrelated
on-campus only:        strong negative pattern appears

filtering on the collider → false link between its causes

The intuition: inside the gate, the two causes must trade off. To pass with weak academics you needed strong sports, and the reverse. The gate converts "either is enough" into "one predicts the absence of the other".

A real example you have seen

Why do the most attractive actors seem to be the weakest actors? Actors get hired for looks or talent — hiring is a collider. Among the hired, the two trade off, so a real inverse pattern appears on screen without existing in the world. The same gate-logic explains "great restaurants have rude staff" and "the friendliest shops are pricey". Anything that survives on either charm or quality shows the trade-off inside.

Remember this

  • A collider has two arrows pointing into it — an effect of both variables.
  • Filtering or adjusting on it invents a link between its causes.
  • Every dataset with an entry gate — hired, admitted, survived, published — is pre-filtered through a collider.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Verified with numpy 1.26.4.

Manufacturing a correlation out of nothing

Two abilities, generated independently — correlation zero by construction. One admission gate. Watch the correlation appear.

collider.py
import numpy as np

rng = np.random.default_rng(8)
n = 100_000
studies = rng.normal(0, 1, n)
sports = rng.normal(0, 1, n)               # generated independently: no link at all
admitted = studies + sports > 1.5          # college admits on the combined score

print(f"all applicants:     corr = {np.corrcoef(studies, sports)[0, 1]:+.3f}")
print(f"admitted students:  corr = {np.corrcoef(studies[admitted], sports[admitted])[0, 1]:+.3f}")
print(f"admission rate: {admitted.mean():.1%}")
Output
all applicants:     corr = +0.006
admitted students:  corr = -0.676
admission rate: 14.5%

The walkthrough

From +0.006 to -0.676 with one boolean mask. No data was edited, no noise added. Selecting rows through a collider created a strong negative correlation between two independent variables. Run any model on the admitted subset and it will confidently learn a relationship that exists nowhere outside the gate.

The mechanism is the thresholded sum. Knowing an admitted student has weak studies tells you their sports must have been strong — the sum had to clear 1.5. Information flows between the causes through the conditioned collider. In DAG terms: conditioning on a collider unblocks the path between its parents.

Selection bias is this exact effect at dataset scale. Your churn dataset contains only signed-up users. Your loan-default data contains only approved applicants. Your "what makes papers cited" corpus contains only published papers. Each gate was a collider, and relationships measured inside are bent — sometimes reversed — relative to the world.

Adjusting is as dangerous as filtering. Adding the collider as a regression control conditions on it numerically, producing the same phantom link. "Control for everything measurable" dies here for good.

Common mistakes

Studying survivors and concluding about entrants. WW2 engineers famously mapped bullet holes on returning aircraft; the holes marked the survivable spots. Armouring where returning planes were hit protects exactly the wrong places. Ask what gate your rows passed through before trusting any pattern in them.

Building models on post-gate data, deploying pre-gate. A credit model trained on approved loans meets the full applicant pool in production — a population whose patterns the gate distorted, connecting to data leakage territory: the training distribution is not the deployment distribution.

Conditioning on a collider's descendant. Filtering on anything downstream of the collider (campus GPA, which follows admission) leaks the same bias. The rule from the DAG lesson includes descendants for exactly this reason.

Assuming small selection means small bias. The admission rate above is 14.5%, and the phantom correlation is -0.68. Harsher gates bend harder — but even mild gates bend.

Try it yourself

Soften the gate: admit with probability rising in studies + sports (a sigmoid) instead of a hard threshold, and watch the phantom correlation weaken but persist. Then build the M-shape: make two hidden causes drive (studies, admission) and (sports, admission) respectively, condition on admission, and check the correlation between the hidden causes — collider bias reaching variables never measured together.

What to learn next

Researcher — Mathematics and papers.

The graphical rule

For a collider $C$ with parents $A \to C \leftarrow B$: the path $A \to C \leftarrow B$ is blocked when $C$ (and all its descendants) are outside the conditioning set, and unblocked when $C$ or any descendant is conditioned on — the collider clause of d-separation from the DAG lesson. Conditioning includes sample selection: analysing only $C = 1$ rows is conditioning on $C$.

For jointly normal independent $A, B$ with $C^* = A + B$ and selection $C^* > c$, the truncated correlation is negative with magnitude growing in the selection severity — the simulation's -0.676 at a 14.5% pass rate. Berkson (1946) derived the hospital-admissions version: two diseases independent in the population become negatively associated among in-patients when either suffices for admission.

Selection bias as collider stratification

Hernán, Hernández-Díaz and Robins (2004), A structural approach to selection bias (Epidemiology), unify case-control selection, loss to follow-up, healthy-worker bias and volunteer bias as conditioning on a common effect $S$ (selection into the analysed sample). The structure covers:

  • M-bias: conditioning on a pre-treatment collider $M$ in an M-shaped graph opens a path between latent causes of $T$ and $Y$ — a pre-treatment variable that harms when controlled, sharpening "adjust for all pre-treatment covariates" into graph-based selection (Ding and Miratrix, 2015, J. Causal Inference).
  • Attrition: dropout caused by both treatment and prognosis makes complete-case analysis collider-conditioned; inverse probability of censoring weights are the standard repair.
  • Sample ratio mismatch in experiments: differential logging survival is selection on a treatment-affected $S$, voiding randomisation's guarantee.

Recoverability

Selection bias is not always fatal: with selection indicator $S$ and a set $W$ separating $(T, Y)$ from $S$ appropriately, conditional effects are recoverable from selected data, sometimes combined with external population data. Bareinboim and Pearl (2012, AAAI; 2016, PNAS Data fusion) give complete graphical conditions. Without such structure, bounds (Manski-style) are the honest output.

ML-specific manifestations

  • Training on gated data (approved loans, admitted patients) then scoring the ungated population — reject inference in credit scoring is a 40-year-old literature on exactly this.
  • Feedback loops: a deployed model influences which rows get labelled next (recommendations, policing data) — collider structure in time.
  • Benchmark curation: dataset inclusion criteria correlated with both features and labels bend measured relationships; Torralba and Efros (2011) document cross-dataset generalisation failures consistent with this.

Reading

  • Berkson (1946), Limitations of the application of fourfold table analysis to hospital data, Biometrics Bulletin.
  • Hernán, Hernández-Díaz and Robins (2004) — the structural taxonomy.
  • Elwert and Winship (2014), Endogenous selection bias: the problem of conditioning on a collider variable, Annual Review of Sociology — the accessible survey.

What to learn next