Baselines and Choosing a Model
Build the whole pipeline with a fake model
Wire data loading, training, evaluation, saving, and serving around a dummy model first — then improving the model becomes a one-line change inside a working system.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Build the entire pipeline — load, train, evaluate, save, predict — around a deliberately dumb model first, and only then improve the model.
Before a wedding, the hall runs a full rehearsal with a stand-in couple. Lights, music, food timing, seating — every system gets tested. Nobody waits for the real couple to discover the microphone is dead.
The dumb model is your stand-in couple. It exercises every pipe in the system while being replaceable in one line the moment the real star arrives.
Why it exists
The model is one box in a chain of boxes: read the data, split it, train, measure, save the result, answer live questions. In practice, most project failures live in the other boxes — a broken join, a mismatched column order, an evaluation reading the wrong file.
If you perfect the model first, every plumbing failure appears after weeks of investment and gets tangled with model debugging. If you build the plumbing first around a fake model, plumbing failures appear on day one, alone, cheap.
How it works
day one:
load → split → train(DUMMY) → evaluate → save → predict one case
✓ ✓ ✓ ✓ ✓ ✓
the whole chain runs, end to end, badly
day two onward:
swap DUMMY for a real model ← a one-line change
every improvement now flows through proven pipesThe skeleton also hands you two gifts. The dummy's score is your first baseline — the number every later model must beat. It is recorded by the same code that will score every future model. And "predict one case" forces the serving question — what arrives, what returns — while the answer is still cheap to change.
A real example you have seen
Film crews shoot with stand-ins to set lighting before the lead actor steps in. New restaurants run trial dinners for friends before opening night. The pattern is old: run the full system with cheap inputs, so real inputs meet a tested system.
Remember this
- Build every pipe around a dummy model first; make the whole chain run on day one.
- The model becomes a one-line swap inside a working system.
- The dummy's score doubles as your recorded baseline, measured by the same code.
What to learn next
- Logistic regression, gradient boosting, or a neural net? — choosing what to swap in.
- What is MLOps? — the discipline the skeleton grows into.
- The baselines you must beat first — the dummy's other job.
Developer — Code and libraries.
Setup
pip install scikit-learnOutputs verified with scikit-learn 1.7 and numpy 1.26.
The skeleton
Every function here survives to production unchanged. Only MODEL is temporary.
import pickle
import numpy as np
from sklearn.datasets import make_classification
from sklearn.dummy import DummyClassifier
from sklearn.metrics import f1_score
from sklearn.model_selection import train_test_split
# The one line you will change later.
MODEL = DummyClassifier(strategy="most_frequent")
def load_data():
X, y = make_classification(n_samples=4000, n_features=10, weights=[0.85],
random_state=0)
return train_test_split(X, y, stratify=y, random_state=0)
def train(X_tr, y_tr):
return MODEL.fit(X_tr, y_tr)
def evaluate(model, X_te, y_te):
return f1_score(y_te, model.predict(X_te))
def save(model, path="model.pkl"):
with open(path, "wb") as f:
pickle.dump(model, f)
def predict_one(path, row):
with open(path, "rb") as f:
return float(pickle.load(f).predict_proba([row])[0, 1])
X_tr, X_te, y_tr, y_te = load_data()
model = train(X_tr, y_tr)
print(f"F1: {evaluate(model, X_te, y_te):.3f}")
save(model)
print(f"one live-style prediction: {predict_one('model.pkl', X_te[0]):.2f}")F1: 0.000 one live-style prediction: 0.00
Now the swap — change two lines in skeleton.py, touch nothing else, and run it again:
from sklearn.ensemble import HistGradientBoostingClassifier
MODEL = HistGradientBoostingClassifier(random_state=0)F1: 0.724 one live-style prediction: 0.00
The walkthrough
F1 of 0.000 is a success. F1 — the balance of precision and recall — is zero because the dummy never predicts the rare class. The pipeline proved it can load, train, score, save, and serve. The badness is the model's, and the model is disposable.
predict_one goes through the saved file, not the in-memory model. This catches serialisation bugs — version drift, unpicklable custom objects — on day one. The saved-artifact path is the one production will use, so the skeleton tests that path from the start.
The swap changed the F1 line and nothing else. That is the property being purchased: every later experiment — new features, new model families, tuning — flows through pipes that already carried a model to a saved file and back.
Real projects grow each function, not the structure. load_data learns to read your warehouse; evaluate gains slice metrics; save gains versioning. The shape survives; structuring the repo around it comes next.
Common mistakes
Polishing the model inside a notebook while the pipeline stays imaginary. The famous last words: "deployment is the easy part, we'll do it at the end". Serving constraints discovered late — input format, latency, missing features at prediction time — invalidate finished models.
A skeleton that skips saving and loading. Training and evaluating in one process hides the artifact boundary, where a large share of real bugs live: pickle protocol changes, library version drift, objects that cannot serialise.
Letting the dummy linger. The skeleton is scaffolding for the first week, not a shrine. Once real models flow, the dummy's remaining job is regression testing — if the pipeline ever scores the dummy above the real model, the pipeline broke.
Skipping the skeleton because "it's a small project". The skeleton is small — the version above is forty lines. Small projects skip it and become the medium-sized messes that needed it.
Try it yourself
Break the pipeline deliberately: make predict_one pass the row's features in reversed order. The saved model still answers — with quietly wrong numbers, and no error raised. Now add a check that catches it (hint: save the expected feature count, or use a DataFrame with named columns). This silent failure is the single most instructive bug in applied ML.
What to learn next
- Logistic regression, gradient boosting, or a neural net? — choosing what to swap in.
- What is MLOps? — the discipline the skeleton grows into.
- The baselines you must beat first — the dummy's other job.
Researcher — Mathematics and papers.
The pattern's provenance
"Walking skeleton" is a general software-engineering pattern — a minimal end-to-end implementation validating the architecture before features arrive (Cockburn; popularised in Growing Object-Oriented Software, Guided by Tests, Freeman & Pryce, 2009). Its ML instantiation appears explicitly in Zinkevich's Rules of Machine Learning (2016): Rule 4, "keep the first model simple and get the infrastructure right", and Rule 13's emphasis that launch plumbing dominates early effort.
The empirical justification is the technical-debt literature: Sculley et al. (2015), Hidden Technical Debt in Machine Learning Systems (NeurIPS), estimate the model as a small fraction of a production system's code and identify the surrounding infrastructure — data dependencies, configuration, serving — as the dominant failure surface. The skeleton front-loads the validation of exactly that surface.
Interface stability as an experimental control
Fixing the pipeline while varying only MODEL implements a controlled comparison: every candidate is trained, evaluated, and serialised by identical code, eliminating pipeline-variance confounds from model comparisons. This matters because evaluation-code heterogeneity is a documented source of irreproducibility — surveyed across fields in Kapoor & Narayanan (2023), Leakage and the Reproducibility Crisis in ML-based Science (Patterns), where inconsistent preprocessing between compared models is a recurring category.
The sklearn estimator API — fit / predict / predict_proba as a uniform contract (Buitinck et al., 2013, API design for machine learning software) — is what makes the one-line swap possible; the pattern generalises to any framework with a stable train/infer interface.
The serving path as part of the hypothesis
predict_one through the serialised artifact tests the deployment function, not only the model: serialisation round-trip, feature ordering, dtype coercion. Training-serving skew — divergence between the training-time and serving-time transformation of the same logical input — is a leading production failure mode (Breck et al., 2017, The ML Test Score, IEEE Big Data, which allocates a full test category to it). A skeleton that exercises the serve path from day one converts skew from a deployment surprise into a unit-testable property.
Limits
The skeleton validates mechanics, not statistics: it cannot detect leakage, weak labels, or an unachievable target — those need the feasibility probe and the audit. The two practices are complements: the probe validates the task, the skeleton validates the machine around it.
What to learn next
- Logistic regression, gradient boosting, or a neural net? — choosing what to swap in.
- What is MLOps? — the discipline the skeleton grows into.
- The baselines you must beat first — the dummy's other job.