Baselines and Choosing a Model
The baselines you must beat first
Every model's score is meaningless until compared against the majority guess, a random guess, and the rule a human would write — three baselines that cost minutes.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A baseline is the score of the dumbest reasonable approach, and no model result means anything until it beats one.
Before praising a new shortcut to the market, time the old road. "Twenty minutes" means nothing on its own — the old road might take nineteen. Every measurement needs the boring comparison point, or it is a number floating in space.
In machine learning that comparison point is the baseline: what the laziest sensible method scores on your exact data.
Why it exists
Model scores flatter shamelessly out of context. "94% accurate" sounds like a triumph. Then you learn that 94% of the emails were not spam. Guessing "not spam" every time scores the same 94% while catching nothing.
Baselines exist to catch this before your announcement does. They cost minutes and have embarrassed every practitioner at least once. The habit of running them first is what the embarrassment teaches.
How it works
Three baselines, always, in this order:
1. MAJORITY predict the most common answer, every time
→ the score gravity gives you for free
2. RANDOM guess in proportion to how common each answer is
→ what luck alone achieves
3. THE RULE the one-line rule a person would write
→ what human common sense achievesEach model you build then has one honest question to answer: which baselines does it beat, and by how much? A model losing to the majority guess is broken. A model barely beating the human rule must justify its complexity in other ways — speed, coverage, maintenance.
A real example you have seen
"Tomorrow's weather equals today's" is a famous forecasting baseline, and it is genuinely hard to beat one day ahead. Weather services measure themselves against it. When a forecaster claims skill, the claim means: better than assuming nothing changes.
Remember this
- A score without a baseline is a number floating in space.
- The three: majority guess, proportional random guess, the one-line human rule.
- Run them before the model, not after — they define what "good" means here.
What to learn next
- Build the whole pipeline with a fake model — the dummy model earns a second job.
- A comparison that actually proves something — when "beats the baseline" needs statistical teeth.
- Model evaluation — the metrics the baseline table is built from.
Developer — Code and libraries.
Setup
pip install scikit-learnOutputs verified with scikit-learn 1.7 and numpy 1.26.
Three baselines and a model, one table
Churn data where 10% of customers leave. Watch which columns expose which pretenders.
import numpy as np
from sklearn.datasets import make_classification
from sklearn.dummy import DummyClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import recall_score
from sklearn.model_selection import train_test_split
# Churn-shaped data: only 10% of customers actually leave.
X, y = make_classification(n_samples=6000, n_features=12, n_informative=6,
weights=[0.9], flip_y=0.03, random_state=2)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, stratify=y, random_state=0)
models = {
"always 'stays'": DummyClassifier(strategy="most_frequent"),
"random guessing": DummyClassifier(strategy="stratified", random_state=0),
"logistic reg": LogisticRegression(max_iter=1000),
}
print(f"{'model':16s} accuracy recall on leavers")
for name, m in models.items():
m.fit(X_tr, y_tr)
pred = m.predict(X_te)
print(f"{name:16s} {(pred == y_te).mean():8.3f} {recall_score(y_te, pred):17.3f}")model accuracy recall on leavers always 'stays' 0.889 0.000 random guessing 0.783 0.102 logistic reg 0.916 0.347
The walkthrough
DummyClassifier is the point of this lesson. It never looks at the features — most_frequent answers the majority class forever, stratified guesses at the base rates. sklearn ships it precisely so baselines cost two lines.
Read the accuracy column and despair correctly. The do-nothing baseline scores 0.889. The trained model scores 0.916 — an apparent lead of 2.7 points over a model that learned nothing. Anyone celebrating 91.6% here is celebrating mostly gravity.
Read the recall column and recover. The baseline catches 0% of leavers; the model catches 34.7%. That is the model's actual contribution, invisible in accuracy. Choosing the column that reveals contribution is the metric lesson's whole argument.
The third baseline is missing deliberately. The one-line human rule ("flag customers inactive 30+ days") needs domain columns this synthetic data lacks. On real data it is often the hardest of the three to beat — write it, always.
Common mistakes
Running baselines after the model. Backwards psychology: the baseline becomes an enemy to defeat instead of a measuring stick. Baseline numbers belong in the project brief before modelling starts.
Beating the majority guess and stopping. The majority guess is the floor, not the bar. The bar is the human rule, plus the cost of ML's complexity. A model 1% better than a cron job with an if-statement should usually lose the argument.
Baseline on one metric, model on another. Compare every row of the table on every metric you report. A model can beat the baseline on accuracy while losing on the metric the business feels.
Forgetting regression has baselines too. Predicting the mean, predicting the median, predicting "same as last time" — DummyRegressor covers the first two. Time-series work leans on the last one; see forecast evaluation.
Try it yourself
Add a fourth row: DummyClassifier(strategy="constant", constant=1) — always predict "leaves". Before running, predict its accuracy and recall from the class balance alone. Then check. Being able to predict a dummy's score from the base rate means you understand both.
What to learn next
- Build the whole pipeline with a fake model — the dummy model earns a second job.
- A comparison that actually proves something — when "beats the baseline" needs statistical teeth.
- Model evaluation — the metrics the baseline table is built from.
Researcher — Mathematics and papers.
What a baseline estimates
The majority-class baseline realises the error of the best constant predictor:
$$ \varepsilon_{\text{const}} = 1 - \max_k P(Y = k) $$
Where $P(Y=k)$ is the marginal class distribution. The gap $\varepsilon_{\text{const}} - \varepsilon_{\text{model}}$ is an estimate of the usable mutual information between features and label that the model family managed to extract — the quantity every project claim rests on. Reporting model error without $\varepsilon_{\text{const}}$ removes the reader's ability to compute the gap; skill scores make it explicit:
$$ \text{skill} = 1 - \frac{\varepsilon_{\text{model}}}{\varepsilon_{\text{ref}}} $$
with $\varepsilon_{\text{ref}}$ the reference (baseline) error. Meteorology has used skill scores against climatology and persistence baselines since Brier (1950); ML reporting is, in this respect, behind a 1950 standard.
Baselines as degenerate hypothesis classes
Each baseline is the optimum of a nested hypothesis class: constants ⊂ univariate thresholds (the human rule) ⊂ full models. Comparing along the chain is a manual structural risk minimisation sweep (Vapnik, 1995): each step up in capacity must pay for itself in validation performance. The chain also localises failure — a full model losing to a threshold rule indicts the optimisation or features, not the task.
Documented baseline upsets
Strong baselines routinely embarrass published methods, which is why they are epistemically load-bearing:
- Dacrema, Cremonesi, Jannach (2019), Are We Really Making Much Progress? (RecSys): of 18 neural recommender papers, only 7 were reproducible, and 6 of those were beaten by tuned nearest-neighbour or graph baselines.
- Makridakis et al. (2018), M4 competition analysis: pure ML methods underperformed statistical baselines and hybrids on large-scale forecasting.
- Grinsztajn et al. (2022), Why do tree-based models still outperform deep learning on tabular data? (NeurIPS D&B): tuned gradient boosting as the persistent tabular baseline.
The pattern generalises: an untuned baseline is a strawman. Baseline tuning effort should be stated — Sculley et al. (2018), Winner's Curse? (ICLR workshop), argue unfair baseline treatment inflates a substantial fraction of reported gains.
Beyond constants: null-model significance
Whether a model beats a baseline is a paired statistical question on shared evaluation data. Permutation tests against label-shuffled data bound the "no-signal" distribution (sklearn.model_selection.permutation_test_score); paired-fold comparisons handle model-versus-baseline — mechanics in a comparison that actually proves something.
What to learn next
- Build the whole pipeline with a fake model — the dummy model earns a second job.
- A comparison that actually proves something — when "beats the baseline" needs statistical teeth.
- Model evaluation — the metrics the baseline table is built from.