Causal DAGs and what to control for
A causal DAG is an arrow diagram of what pushes what — and reading it tells you which variables to adjust for, and which ones adjusting would ruin.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A causal DAG is a diagram of arrows showing which things push on which — a wiring map of the system you are studying.
Think of a family tree. Arrows run one way — parents to children — and nobody is their own ancestor. A causal diagram works the same way: an arrow from "exercise" to "health" means exercise pushes on health. DAG stands for directed acyclic graph — arrows have direction, and no loop lets a thing cause itself.
The diagram is not decoration. Reading it mechanically answers the hardest practical question in observational work: which variables should I adjust for?
Why it exists
The confounding lesson ended with an instinct: adjust for common causes. But real systems tangle. Is "resting heart rate" a confounder of exercise and health, or a result of exercise? Adjusting feels safe either way. It is not.
Some variables sit on the path from cause to effect. Such a variable is a mediator: a middleman carrying the effect onward. Adjust for a mediator and you erase part of the very effect you are measuring. The DAG makes the difference visible at a glance: confounders stand behind the cause; mediators stand between cause and effect.
How it works
fitness ← confounder: arrows OUT to both
↙ ↘ ADJUST for it
exercise → health
↘ ↗
heart rate ← mediator: sits ON the path
LEAVE it aloneHere is the rule of thumb, readable from any DAG. Block every backdoor path: any route from cause to effect that starts with an arrow into the cause. The fitness route above is one such backdoor. Adjusting for fitness closes it. The heart-rate route is not a backdoor — it is part of the effect's own delivery road. Adjusting for it closes the road you are measuring.
A real example you have seen
"Does education raise income?" Family wealth pushes on both education and income — a backdoor, adjust for it. But "occupation" sits between education and income: education changes what job you get, and the job pays. Studies that adjust for occupation report education "barely matters" — after quietly removing the main way it matters. The DAG catches this in seconds.
Remember this
- A DAG draws what pushes on what; no cycles allowed.
- Confounders (arrows out to both sides) — adjust. Mediators (on the path) — leave alone.
- The mechanical rule: close every backdoor path, touch nothing on the front road.
What to learn next
- Colliders and selection bias — the third variable type, where controlling creates bias.
- Simpson's paradox — DAG reading turning a paradox into a bookkeeping exercise.
- Confounding — the backdoor story this lesson mechanised.
Developer — Code and libraries.
Setup
pip install numpy scikit-learnVerified with numpy 1.26.4 and scikit-learn 1.7.2.
Three adjustment choices, three different answers
We build the exact DAG from the beginner block, with known arrow strengths, and try the three possible control sets.
import numpy as np
from sklearn.linear_model import LinearRegression
rng = np.random.default_rng(6)
n = 50_000
fitness = rng.normal(0, 1, n) # confounder
exercise = 0.8 * fitness + rng.normal(0, 1, n) # fitness -> exercise
heart = -0.7 * exercise + rng.normal(0, 1, n) # mediator: exercise lowers resting HR
health = 0.5 * exercise + 0.9 * fitness - 0.4 * heart + rng.normal(0, 1, n)
# total true effect of exercise: 0.5 direct + (-0.7 x -0.4) via heart rate = 0.78
def effect(controls):
X = np.column_stack([exercise] + controls)
return LinearRegression().fit(X, health).coef_[0]
print(f"control nothing: {effect([]):+.3f} <- backdoor open, biased up")
print(f"control fitness: {effect([fitness]):+.3f} <- correct (truth +0.780)")
print(f"control fitness + heart rate: {effect([fitness, heart]):+.3f} <- blocked our own path")control nothing: +1.220 <- backdoor open, biased up control fitness: +0.775 <- correct (truth +0.780) control fitness + heart rate: +0.500 <- blocked our own path
The walkthrough
Three defensible-looking analyses, three answers: +1.22, +0.78, +0.50. Same data, same model class. Only the control set differs. Without a DAG, choosing between them is folklore; with one, it is mechanical.
Controlling nothing leaves the backdoor open. Fitness pushes exercise up and health up, adding a fake +0.44 to the true +0.78.
Controlling fitness alone is exactly right. The backdoor closes; the front road through heart rate stays open, so the total effect — direct plus mediated — comes through: +0.775 against a truth of +0.78.
Adding heart rate subtracts the mediated path. The estimate collapses to the direct effect (+0.50), silently answering a different question: "what would exercise do if we clamped everyone's heart rate?" Sometimes that is the question — mediation analysis asks it deliberately. Answering it by accident is the error.
The regression is doing the adjusting. Each control variable added to the design matrix is the smooth version of stratifying on it — the connection made in confounding.
Common mistakes
"More controls are always safer." The output above is the counterexample. Mediators shrink the estimate; the next lesson's colliders corrupt it outright. Control sets are chosen by graph position, never by availability.
Controlling for post-treatment variables. Anything measured after the treatment could sit downstream of it — which makes it a mediator or collider risk. Pre-treatment variables are the safe hunting ground for confounders.
Drawing the DAG to fit the desired answer. The DAG encodes assumptions, and assumptions can be wrong. Draw it before analysing; expose it to colleagues who know the domain; test the independences it implies where possible.
Believing the DAG must be complete to be useful. Even a partial DAG rules control sets in and out. "I do not know whether A causes B" can itself be drawn (try both, report both).
Try it yourself
Add a variable motivation with arrows into exercise and into health, but leave it unmeasured — generate it, use it in the data, then refuse to control for it. Watch every control set miss the truth. That gap is unmeasured confounding as the DAG displays it. Then measure it and confirm [fitness, motivation] restores +0.78.
What to learn next
- Colliders and selection bias — the third variable type, where controlling creates bias.
- Simpson's paradox — DAG reading turning a paradox into a bookkeeping exercise.
- Confounding — the backdoor story this lesson mechanised.
Researcher — Mathematics and papers.
d-separation and the backdoor criterion
A path between $T$ and $Y$ is blocked by conditioning set $Z$ iff it contains a chain $\cdot \to m \to \cdot$ or fork $\cdot \leftarrow m \to \cdot$ with $m \in Z$, or a collider $\cdot \to c \leftarrow \cdot$ with $c \notin Z$ and no descendant of $c$ in $Z$. Two nodes are d-separated by $Z$ when every path between them is blocked; d-separation implies conditional independence in every distribution the DAG can generate (Pearl, 1988).
Backdoor criterion (Pearl, 1993): $Z$ suffices to identify the effect of $T$ on $Y$ if (i) no node in $Z$ is a descendant of $T$, and (ii) $Z$ blocks every path from $T$ to $Y$ that begins with an arrow into $T$. Then
$$ P(y \mid do(t)) = \sum_z P(y \mid t, z)\, P(z) $$
— the adjustment formula, with all symbols as in the confounding lesson's g-formula. In the simulation, $Z = {\text{fitness}}$ satisfies the criterion; $Z = {\text{fitness}, \text{heart}}$ violates condition (i), and the estimate duly shifts from total to controlled-direct effect.
Beyond backdoor
- Front-door criterion (Pearl, 1995): identification through a fully-observed mediator even with unmeasured $T$–$Y$ confounding — rare in practice, conceptually important.
- do-calculus (Pearl, 1995): three rewrite rules, complete for nonparametric identification (Shpitser and Pearl, 2006; Huang and Valtorta, 2006). The ID algorithm decides identifiability of any $P(y \mid do(t))$ from a DAG with latent variables.
- Adjustment-set selection: among valid sets, statistical efficiency differs — prefer outcome-parents over treatment-parents (Henckel, Perković and Maathuis, 2022, JRSS-B). Software:
dagitty(Textor et al.) enumerates valid sets from a drawn DAG.
Effects vocabulary
With mediator $M$: total effect = natural direct + natural indirect (VanderWeele, 2015, Explanation in Causal Inference). The developer block's third regression estimates the controlled direct effect at fixed $M$ under linearity — distinct from the natural direct effect in nonlinear systems. Mediation formulae require cross-world assumptions beyond the backdoor set; treat casual "we controlled for the mechanism" claims with suspicion.
Testability and discovery
A DAG implies a set of conditional independences (its Markov properties); these are testable, though multiple DAGs share them (Markov equivalence). Structure-discovery algorithms (PC, GES, and successors — Spirtes, Glymour and Scheines, 2000, Causation, Prediction, and Search) recover equivalence classes from data under strong assumptions. In applied work the DAG is usually asserted from domain knowledge and stress-tested, not discovered.
Reading
- Pearl (2009), Causality, 2nd ed. — chapters 1 and 3.
- Cinelli, Forney and Pearl (2022), A crash course in good and bad controls, Sociological Methods & Research — the definitive control-selection cheat sheet.
- Hernán and Robins (2020), Causal Inference: What If, chapters 6–8.
What to learn next
- Colliders and selection bias — the third variable type, where controlling creates bias.
- Simpson's paradox — DAG reading turning a paradox into a bookkeeping exercise.
- Confounding — the backdoor story this lesson mechanised.