Linear Models and Regularisation
Poisson regression for counts
Counts are non-negative, whole, and grow multiplicatively — Poisson regression models them that way, where a straight line predicts nonsense.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Poisson regression predicts counts — 0, 1, 2, 3 — and treats every influence as a multiplier instead of an add-on.
Watch customers arrive at a chai stall for an hour. Some hours bring 2, some 5, one strange hour brings 11. Arrivals come one by one, at random moments, and the count is always a whole number, never negative. That arrival pattern has a name: Poisson. It describes a startling amount of the world. Calls to a helpline, potholes per kilometre, goals per match, typos per page.
Poisson regression predicts such counts from features: how many customers this hour, given the weather, the day, the adverts running.
Why it exists
Try ordinary linear regression on counts and two things go wrong.
First, the range. A straight line happily predicts -1.4 customers for a rainy Monday. There is no such thing as minus customers.
Second, the way influences combine. A festival day does not add three customers to every stall equally. It adds thirty to a busy stall and one to a quiet one — influences on counts tend to multiply. "Festival doubles traffic" is the natural sentence; "festival adds 4.2 visits" is not.
Poisson regression fixes both with one move: the model works with multipliers, and its predictions can approach zero but never cross it. It is the counts attachment of the GLM machine.
How it works
linear: visits = base + this + that (can go negative)
poisson: visits = base x this x that (always positive)
each advert x1.5 festival x2 rain x0.6
quiet stall: 2 x 1.5 = 3 busy stall: 20 x 1.5 = 30A real example you have seen
Insurance premiums: claims per driver per year is a count, priced by multipliers — young driver ×1.8, city ×1.3. Epidemiology counts cases per district per week the same way. Football analytics models goals per match as Poisson, which is how betting odds for exact scores are set.
Remember this
- Counts are whole and non-negative; a straight line ignores both facts.
- Poisson regression's influences multiply — "doubles the rate", not "adds 4.2".
- It is the counts attachment of the GLM machine.
What to learn next
- Generalised linear models — the machine this attachment clicks into.
- Quantile regression — when you need the busy-day tail, not the average.
- Time series basics — counts over time bring their own structure.
Developer — Code and libraries.
Setup
pip install scikit-learnOutputs verified with scikit-learn 1.7.2 on CPU.
Adverts and shop visits
import numpy as np
from sklearn.linear_model import LinearRegression, PoissonRegressor
rng = np.random.default_rng(3)
ads = rng.integers(0, 6, 60).astype(float) # adverts run each week
visits = rng.poisson(np.exp(0.8 + 0.35 * ads)) # customer visits that week
X = ads.reshape(-1, 1)
lin = LinearRegression().fit(X, visits)
poi = PoissonRegressor(alpha=0).fit(X, visits)
print(f"linear : visits = {lin.intercept_:.1f} + {lin.coef_[0]:.1f} * ads")
print(f"poisson: each advert multiplies visits by {np.exp(poi.coef_[0]):.2f}")
for n in (0, 5, 10):
a = np.array([[float(n)]])
print(f"ads={n:2d}: linear {lin.predict(a)[0]:6.1f} poisson {poi.predict(a)[0]:6.1f}")linear : visits = 0.8 + 2.5 * ads poisson: each advert multiplies visits by 1.49 ads= 0: linear 0.8 poisson 2.1 ads= 5: linear 13.2 poisson 15.2 ads=10: linear 25.7 poisson 111.8
The walkthrough
The models tell different stories. Linear: "each advert adds 2.5 visits, whatever the current level." Poisson: "each advert multiplies visits by 1.49." The data was generated multiplicatively (each advert multiplies the rate by exp(0.35) ≈ 1.42), and only one model can even say that sentence.
At zero adverts, linear predicts 0.8 visits — beneath the true quiet-week average of about 2.2, dragged low by forcing one straight line through curved data. Add a feature with a negative effect (rain, roadworks) and linear regression will cross into negative visits without blinking.
At ten adverts, the models disagree by 4x (25.7 vs 111.8). Both are extrapolating beyond the training range of 0–5, so treat both numbers with suspicion — but note the shapes: exponential growth versus a straight line is a modelling decision with enormous consequences outside the data. Exponential extrapolation can also overshoot badly; honesty about the range matters in both directions.
np.exp(poi.coef_[0]) is the reading ritual. Poisson coefficients live in log-space; exponentiate to get multipliers. A coefficient of -0.5 means exp(-0.5) ≈ 0.61 — that feature cuts the rate by 39%.
alpha=0 switches off PoissonRegressor's default ridge penalty (alpha=1.0) so the comparison with plain OLS is fair. In production, keep some regularisation.
Common mistakes
Negative or fractional targets. PoissonRegressor refuses negative targets outright — ValueError: Some value(s) of y are out of the valid range of the loss 'HalfPoissonLoss'. Rates like "visits per opening hour" (fractional) belong in the exposure pattern below, not in the target.
Ignoring exposure. Comparing weekly counts across shops open different hours is comparing apples to weeks. Model counts with the time-open as an offset — in statsmodels, sm.GLM(..., family=sm.families.Poisson(), offset=np.log(hours_open)) — so coefficients describe rates, not raw totals.
Trusting Poisson's variance promise. The Poisson family assumes wobble equals the mean: an average of 5 implies variance 5. Real counts are usually messier — one viral post, one festival — a state called overdispersion. Your rate estimates stay reasonable, but the error bars become too confident. Check, and escalate to negative binomial if needed (researcher block).
Log-transforming the target instead. log(visits + 1) into OLS looks similar and behaves worse: the +1 is arbitrary, zeros distort, and predictions are biased on the way back. Poisson regression is the correct version of that instinct.
Try it yourself
Generate a second dataset with visits = rng.poisson(np.exp(0.8 + 0.35 * ads)) * rng.integers(1, 4, 60) — bursty, overdispersed arrivals — and refit. Compare the mean and variance of visits at each ads level to see the broken promise. Do the rate estimates survive?
What to learn next
- Generalised linear models — the machine this attachment clicks into.
- Quantile regression — when you need the busy-day tail, not the average.
- Time series basics — counts over time bring their own structure.
Researcher — Mathematics and papers.
Model
Counts $y_i \in {0, 1, 2, \dots}$ with the log link:
$$ y_i \sim \text{Poisson}(\mu_i), \qquad \log \mu_i = x_i^\top \beta \; (+ \log t_i) $$
Where:
- $\mu_i$ — the expected count; the log link keeps $\mu_i = e^{x_i^\top \beta} > 0$.
- $\beta_j$ — log rate ratios: one unit of $x_j$ multiplies the rate by $e^{\beta_j}$.
- $\log t_i$ — the optional offset for exposure $t_i$ (hours open, person-years), entering with coefficient fixed at 1.
The Poisson pmf $P(y) = e^{-\mu}\mu^y / y!$ has the defining property $\operatorname{Var}(y) = \mathbb{E}[y] = \mu$ — the equidispersion assumption. The log-likelihood is concave under the canonical log link; IRLS/Newton converges globally at $O(nd^2)$ per iteration.
Overdispersion and its fixes
Empirical counts almost always have $\operatorname{Var} > \mu$ (unobserved heterogeneity, contagion, clustering). Consequences: $\hat\beta$ stays consistent (the mean model can be right), but likelihood-based standard errors are too small — inflated significance. Remedies, in escalating order:
- Quasi-Poisson: estimate a dispersion factor $\hat\phi = \frac{1}{n-d}\sum_i (y_i - \hat\mu_i)^2/\hat\mu_i$ and scale standard errors by $\sqrt{\hat\phi}$ (Wedderburn, 1974).
- Negative binomial: a Gamma-mixed Poisson with $\operatorname{Var} = \mu + \alpha\mu^2$; likelihood-based, extra parameter $\alpha$ (Cameron and Trivedi, 1998, Regression Analysis of Count Data — the field's reference).
- Zero-inflated / hurdle models: a separate process for structural zeros (weeks the shop was shut vs. open-but-empty) (Lambert, 1992).
- Robust (sandwich) standard errors as a model-agnostic backstop.
Model choice diagnostics: Pearson dispersion statistic, rootograms, and score tests for zero inflation.
Connections
- Poisson regression with categorical predictors reproduces log-linear contingency-table analysis exactly.
- The Poisson trick: piecewise-exponential survival models fit as Poisson GLMs with time-interval offsets (Holford, 1980).
- In ML terms, minimising Poisson deviance $2\sum_i [y_i \log(y_i/\hat\mu_i) - (y_i - \hat\mu_i)]$ is the count analogue of squared error; LightGBM and XGBoost expose
objective="poisson", and neural recommenders use log-link Poisson output layers for interaction counts. - Insurance pricing GLMs (frequency × severity, Poisson × Gamma) remain the regulated industry standard, with the Tweedie family bridging the two (Jørgensen, 1987).
What to learn next
- Generalised linear models — the machine this attachment clicks into.
- Quantile regression — when you need the busy-day tail, not the average.
- Time series basics — counts over time bring their own structure.