Causal Inference Basics

Difference-in-differences

Difference-in-differences measures a change against a untouched comparison group, so shared trends — seasons, economies, fashions — cancel out of the answer.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Difference-in-differences (DiD) measures the effect of a change by comparing your before-and-after jump against the before-and-after jump of someone who did not change anything.

Two tea stalls sit on the same street. One installs a fancy espresso machine in June. Its sales rise 20% by August — but it is also mango season, tourists arrived, and exams ended. The other stall, unchanged, saw sales rise 12% over the same weeks. The street's shared tide lifted both boats by 12. The machine's own contribution is the difference of the differences: 20 minus 12 — 8%.

Two subtractions, hence the name. The first removes each stall's own baseline. The second removes the tide.

Why it exists

Before-and-after comparisons are seductive and broken — time never holds still while you act. Sales after the renovation include the renovation and the festival season and the economy. A raw before/after mixes them inseparably.

DiD's insight: you do not need time to hold still. You need a companion who rode the same tide without making your change. Their journey measures the tide. Subtract it.

How it works

                before    after     change
renovated shop    50   →   71       +21
neighbour shop    48   →   60       +12   ← the tide, measured

effect = 21 - 12 = +9

The whole method leans on one promise, the parallel trends assumption: without the renovation, your shop's sales would have moved like the neighbour's. The promise is about the invisible road not taken, so it can never be fully verified. But it can be supported. Show the two shops moving in step for months before the change. If they danced together before, believing they would have continued is reasonable.

A real example you have seen

Every "state X raised minimum wage, banned plastic bags, launched a scheme — did it work?" debate is a DiD in the making. A neighbouring state plays the companion. The most famous: Card and Krueger compared fast-food jobs in New Jersey before and after its 1992 minimum-wage rise, against next-door Pennsylvania. They found no job loss, upending a textbook prediction. That study reshaped a policy field and later earned a Nobel.

Remember this

  • DiD = your change minus the companion's change over the same window.
  • The companion absorbs seasons, economies and shared shocks.
  • It all rests on parallel trends — support it by showing the pre-period lockstep.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Verified with numpy 1.26.4.

A renovation with a season in the way

Both shops share a strong seasonal wave. Shop A renovates at week 10 with a true effect of +9. Watch before/after alone get fooled, and DiD recover.

did.py
import numpy as np

rng = np.random.default_rng(14)
weeks = np.arange(20)
season = 50 + 8 * np.sin(weeks / 3)              # both shops share the season
shop_a = season + 12 + rng.normal(0, 2, 20)      # A renovates after week 9
shop_b = season + rng.normal(0, 2, 20)           # B never renovates
shop_a[10:] += 9                                 # true renovation effect: +9

after = weeks >= 10
change_a = shop_a[after].mean() - shop_a[~after].mean()
change_b = shop_b[after].mean() - shop_b[~after].mean()

print(f"shop A, before vs after:  {change_a:+.2f}")
print(f"shop B, same comparison:  {change_b:+.2f}   <- the season alone did this")
print(f"difference-in-differences: {change_a - change_b:+.2f}  (truth: +9.00)")
Output
shop A, before vs after:  +0.47
shop B, same comparison:  -9.64   <- the season alone did this
difference-in-differences: +10.11  (truth: +9.00)

The walkthrough

Before/after alone reports +0.47 — it nearly erases a +9 effect. The sine wave happened to be falling through the after-period, and the raw comparison charged that fall against the renovation. In other seeds, the season inflates instead. Either way, time's fingerprints are all over the naive number.

Shop B's -9.64 is the season, measured. B changed nothing; its drop is pure tide. Subtracting it hands back +10.11, within noise of the truth.

Note the constant +12 gap between shops. A is a bigger shop throughout. DiD never needs the shops to be equal — the first differencing removes any fixed gap. What it needs is equal movement: the parallel trends.

As a regression: stack both shops' weeks and fit sales ~ is_A + is_after + is_A*is_after — the interaction coefficient is the DiD estimate, and linear regression machinery brings standard errors along. With many units and periods this becomes the two-way fixed effects workhorse.

Common mistakes

Skipping the pre-trend plot. The single most convincing exhibit is the two lines moving in step before the change. If they diverged already, the companion is measuring a different tide, and the estimate inherits the divergence. Always plot; report the pre-period gap-stability.

Choosing a companion affected by the treatment. If the neighbour's customers defected to your renovated shop, their decline is your effect leaking into the "tide" measurement — interference wearing an observational hat. Pick companions outside the blast radius.

Naive error bars on autocorrelated series. Weekly sales correlate across weeks; treating 20 weeks as 20 independent observations shrinks error bars dishonestly (Bertrand, Duflo and Mullainathan's famous critique). Cluster the errors by unit, or aggregate to one before/after pair per unit.

Treatments that announce themselves. If shop A's customers anticipated the renovation (a pre-opening sale, hoarding), the "before" period is contaminated. Check for effects appearing before the treatment date — their presence flags anticipation or a broken companion.

Try it yourself

Violate parallel trends on purpose: give shop A its own extra trend 0.4 * weeks and rerun — watch DiD absorb the divergence into a wrong answer, and confirm the pre-period gap between shops was already widening (the tell). Then add a third, untouched shop C and estimate the effect using each companion separately — disagreement between companions is itself a diagnostic.

What to learn next

Researcher — Mathematics and papers.

The estimand and identification

With groups $g \in {0, 1}$ (companion, treated) and periods $t \in {0, 1}$ (before, after), the canonical 2×2 DiD is

$$ \hat{\tau}{DiD} = \left(\bar{Y}{1,1} - \bar{Y}{1,0}\right) - \left(\bar{Y}{0,1} - \bar{Y}_{0,0}\right) $$

  • $\bar{Y}_{g,t}$ — mean outcome for group $g$ in period $t$.

Identification of the ATT requires parallel trends: $\mathbb{E}[Y(0){1} - Y(0){0} \mid g{=}1] = \mathbb{E}[Y(0){1} - Y(0){0} \mid g{=}0]$ — the treated group's untreated change equals the companion's. This is weaker than ignorability (levels may differ arbitrarily) but scale-dependent: parallel in levels and parallel in logs are different assumptions, and at most one usually holds. No-anticipation is the companion assumption.

The two-way fixed effects trap

The regression $Y_{it} = \alpha_i + \lambda_t + \tau D_{it} + \varepsilon_{it}$ (unit and time fixed effects) equals canonical DiD in the 2×2 case. With staggered adoption — units treated at different times — the TWFE $\hat{\tau}$ is a weighted average of many 2×2 comparisons, including "forbidden" ones using already-treated units as controls; weights can be negative, and with heterogeneous effects the estimate can carry the wrong sign (Goodman-Bacon, 2021, J. Econometrics — the decomposition; de Chaisemartin and D'Haultfœuille, 2020, AER). Repairs:

  • Callaway and Sant'Anna (2021, J. Econometrics) — group-time ATTs $ATT(g, t)$ aggregated explicitly.
  • Sun and Abraham (2021, J. Econometrics) — interaction-weighted event studies.
  • Borusyak, Jaravel and Spiess (2024, ReStud) — imputation estimators.

Modern applied work reports event-study coefficients (dynamic effects relative to treatment time) with pre-period placebo coefficients as the parallel-trends exhibit; Roth (2022, AER: Insights) cautions that pre-trend tests have low power and conditioning on passing them distorts inference — Rambachan and Roth (2023, ReStud) provide partial-identification robustness instead.

Inference

Serial correlation inflates naive precision drastically (Bertrand, Duflo and Mullainathan, 2004, QJE): cluster at the unit level (or the treatment-assignment level); with few treated clusters, wild cluster bootstrap (Cameron, Gelbach and Miller, 2008) or randomisation inference. With a single treated unit, the synthetic control method (Abadie, Diamond and Hainmueller, 2010, JASA) constructs the companion as a weighted combination of donors — DiD's small-N sibling.

Reading

  • Card and Krueger (1994), Minimum wages and employment, AER — the canonical study.
  • Angrist and Pischke (2009), Mostly Harmless Econometrics, chapter 5.
  • Roth, Sant'Anna, Bilinski and Poe (2023), What's trending in difference-in-differences?, J. Econometrics — the definitive modern survey.

What to learn next