CI/CD for machine learning
A machine that rebuilds, retests and repackages your model on every change — and refuses to ship it when a check fails, no matter who is asking.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
CI/CD is a machine that rebuilds and rechecks everything on every change, and refuses to ship when a check fails.
CI is short for continuous integration — every change gets combined with everyone else's and tested, straight away. CD is continuous delivery — if the tests pass, the result is packaged and sent onward automatically.
The analogy you have already lived
You have walked through a security check at a mall or an airport. Your bag goes on the belt, through the scanner, every single time.
Nobody is exempt. Not the regular visitor, not the manager, not the person who is late. The machine does not get tired at 11 p.m. on a Friday and wave someone through.
That is CI. The value is not that it is clever. The value is that it cannot be talked out of it.
Why it exists
Picture the version without it.
Somebody makes a change on Friday evening. They mean to run the tests. They are in a hurry, and the tests take eleven minutes, and it is a small change.
On Monday the model returns nothing but zeros. Nobody knows which of the nine changes made that weekend caused it. Three people spend a day finding out.
CI removes the decision. Change goes in, checks run, result is green or red. There is nothing to remember and nothing to skip.
What makes this harder than normal software
Ordinary software testing asks one question: does the code do what I said?
Machine learning needs three more questions. Skipping them is why so many pipelines pass every test and still ship a broken model.
Is the data still what we agreed? The right columns, the right types, values in sensible ranges, not suddenly half empty.
Is the model still good enough? Score it on a frozen set of examples with known answers. If it drops below an agreed line, stop.
Does the model still behave sensibly? Some things must always be true. Raising an applicant's income must never lower their approval score. A rule like that catches disasters no accuracy number will.
How it works
you push a change
|
v
install the pinned libraries
|
train the model
|
test the CODE does it run, does it return the right shape
test the DATA right columns, no floods of missing values
test the MODEL score on frozen examples, sanity rules hold
|
all green? ---- no ----> STOP. Nothing ships. Somebody gets told.
|
yes
v
build the sealed box -> deploy to a copy -> deploy for realEverything above the "STOP" line is CI. Everything below it is CD.
Most teams get enormous value from the top half alone, and it is where you should start.
The frozen exam paper
The single most useful thing on this page is small enough to build today.
Take twenty real examples where you know the correct answer. Save them in a file. Check that file into version control and never regenerate it.
Every time the model changes, it sits that same exam. If the score falls below your line, the change does not ship.
It is twenty rows. It catches the wrong data file and the flipped label column. It catches a feature that stopped being computed. It catches a retrain that quietly failed halfway.
Where you have already benefited from this
Every app on your phone that updates without breaking. Every website that changes overnight and still works in the morning.
You never see CI/CD. You see its absence, on the days something ships broken.
The honest part
CI for machine learning is slower and more irritating than CI for ordinary code.
Training takes minutes or hours, not seconds. Model tests can be flaky, because training has randomness in it. A test that passes four times out of five is worse than no test. People learn to re-run it until it goes green.
The cures are real, and they take discipline. Fix your random seeds. Set your threshold with a margin below today's score, not at it. Keep the slow full retrain on a schedule, and run the fast checks on every change.
Remember this
- CI runs the checks automatically, so nobody can skip them in a hurry.
- Machine learning needs tests for data and model, not only code.
- A frozen set of twenty known examples is the cheapest safety net you will ever build.
What to learn next
- Monitoring and model drift — the signal that should trigger a retrain in the first place.
- Docker for ML — the image this pipeline builds when the tests go green.
- Model deployment — the same discipline applied to language models.
Developer — Code and libraries.
Setup
pip install pytest scikit-learn pandas joblibThis lesson continues from model serving. It uses the model.joblib that the train.py there produces. Run that script first, or the fixture below has nothing to load.
The test file
Six tests. Each one catches a failure that has actually shipped somewhere.
import joblib
import numpy as np
import pandas as pd
import pytest
COLUMNS = ["income", "years", "age"]
MIN_ACCURACY = 0.85
# A frozen set of 20 applicants with known outcomes. Checked into git, never
# regenerated. This is the exam paper the model has to keep passing.
GOLDEN = pd.DataFrame([
(77.5, 4.4, 44, 1), (46.0, 9.5, 27, 1), (78.0, 7.9, 60, 1),
(58.6, 8.7, 63, 1), (57.3, 1.7, 28, 0), (21.2, 0.7, 30, 0),
(78.2, 6.0, 39, 1), (5.5, 1.7, 42, 0), (24.0, 7.3, 51, 1),
(37.6, 4.1, 36, 1), (63.5, 5.3, 57, 1), (19.8, 9.4, 48, 1),
(69.7, 5.2, 64, 1), (78.8, 1.1, 29, 1), (17.3, 1.6, 58, 0),
(49.8, 5.5, 29, 1), (5.7, 5.2, 60, 0), (34.0, 6.4, 53, 1),
(8.3, 4.0, 51, 0), (76.7, 6.5, 25, 1),
], columns=COLUMNS + ["repaid"])
@pytest.fixture(scope="module")
def model():
return joblib.load("model.joblib")
def test_model_expects_exactly_the_agreed_columns(model):
assert list(model.feature_names_in_) == COLUMNS
def test_every_score_is_a_usable_probability(model):
p = model.predict_proba(GOLDEN[COLUMNS])[:, 1]
assert np.isfinite(p).all(), "a NaN reached the caller"
assert ((p >= 0.0) & (p <= 1.0)).all()
def test_more_income_never_lowers_the_score(model):
"""A directional expectation. Nothing else changes, only income rises."""
ladder = pd.DataFrame({"income": np.arange(10.0, 80.0, 5.0),
"years": 4.0, "age": 35.0})
p = model.predict_proba(ladder[COLUMNS])[:, 1]
assert np.all(np.diff(p) >= -1e-9), "score fell when income rose"
def test_accuracy_on_the_frozen_set(model):
accuracy = (model.predict(GOLDEN[COLUMNS]) == GOLDEN["repaid"]).mean()
assert accuracy >= MIN_ACCURACY, f"accuracy fell to {accuracy:.3f}"
def test_a_very_strong_applicant_is_approved(model):
strong = pd.DataFrame([{"income": 75.0, "years": 9.0, "age": 40.0}])
assert model.predict(strong[COLUMNS])[0] == 1
def test_missing_values_are_refused_not_guessed(model):
broken = pd.DataFrame([{"income": np.nan, "years": 4.0, "age": 35.0}])
with pytest.raises(ValueError):
model.predict(broken[COLUMNS])pytest test_model.py -q...... [100%] 6 passed in 0.79s
Now break something on purpose
This is the part worth doing rather than reading. Retrain with the label column inverted. That is a real and common bug: an upstream table changes 1 from meaning "repaid" to meaning "defaulted".
# add this line to train.py, replacing the fit call
model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000)).fit(X, 1 - y)..FFF. [100%] =========================== short test summary info =========================== FAILED test_model.py::test_more_income_never_lowers_the_score - AssertionError: score fell when i... FAILED test_model.py::test_accuracy_on_the_frozen_set - AssertionError: accuracy fell to 0.050 FAILED test_model.py::test_a_very_strong_applicant_is_approved - assert 0 == 1 3 failed, 3 passed in 0.83s
Line widths and the truncation with ... depend on your terminal, so those will differ. What matters is which tests moved.
Three different tests caught one bug, and each one describes it differently. Accuracy fell to 0.050. The score fell when income rose. An applicant with every positive signal was rejected. Any one of them alone would have stopped the release.
Notice which tests still passed. The schema was fine. The probabilities were valid numbers between zero and one. A suite that checked only shapes and types would have shown six green ticks. It would have shipped a model that is exactly backwards.
Wiring it to every push
name: model checks
on:
push:
pull_request:
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.11"
cache: pip
# Pinned versions, so a library release cannot change the result
# of a test run without somebody editing a file first.
- name: Install dependencies
run: pip install -r requirements.txt
- name: Train
run: python train.py
- name: Test the model
run: pytest -q
# Only reached when every test above passed.
- name: Build the image
run: docker build -t loan-model:${{ github.sha }} .No output block here. The log is a live web page with timings and collapsible steps, and reproducing it as text would be a fiction.
The important property is the ordering. docker build sits after pytest, and GitHub Actions stops a job at the first failing step. A red test means no image is ever built, so a broken model cannot reach a registry by accident.
The four kinds of ML test, named
Worth knowing the vocabulary, because the literature uses it.
| Kind | Question | Example above |
|---|---|---|
| Schema | Right columns, right types? | test_model_expects_exactly_the_agreed_columns |
| Minimum functionality | Does the easiest possible case work? | test_a_very_strong_applicant_is_approved |
| Directional expectation | Does changing one input move the output the right way? | test_more_income_never_lowers_the_score |
| Invariance | Does a change that should not matter leave the output alone? | not shown — add one for a field the model must ignore |
Ribeiro et al. (2020) formalised these as behavioural testing, in the paper that introduced CheckList. The idea transfers cleanly from language models to tabular ones.
Common mistakes
Setting the threshold at today's exact score. If the model scores 0.87 and you assert >= 0.87, the next harmless retrain fails the build. Set the line where you would genuinely refuse to ship — here, 0.85 against a model scoring 0.95.
Regenerating the golden set when a test fails. This is the ML equivalent of deleting the failing test. If the frozen examples are wrong, fix them deliberately, in their own commit, with the reason written down.
Retraining from scratch on every commit. For anything larger than a toy this makes CI unusably slow. Split it: fast tests against the last committed model on every push, full retrain nightly or on demand.
Unpinned dependencies in CI. Without pins, a library released this morning can turn your build red with no change from you. Worse, it can turn it green when it should be red.
Testing only the model. The serving layer needs tests too. The TestClient checks from model serving run in under a second and catch the validation bugs that models never see.
No path to roll back. Automated deployment without a one-command rollback is automated risk. Keep the previous image tagged and know the command before you need it.
Try it yourself
Add an invariance test. Give the model two applicants whose only difference is a field you believe it should ignore, and assert the scores match.
If your model has no such field, that finding is more valuable than the test. It means every column you collect is influencing decisions, including any you would struggle to defend.
What to learn next
- Monitoring and model drift — the signal that should trigger a retrain in the first place.
- Docker for ML — the image this pipeline builds when the tests go green.
- Model deployment — the same discipline applied to language models.
Researcher — Mathematics and papers.
The rubric
Breck et al. (2017), The ML Test Score, remains the most useful checklist in the field because it is specific enough to argue with. Twenty-eight tests over four axes, each scored 0, 0.5 or 1, with the final score being the minimum across axes — deliberately, so a team cannot compensate for absent monitoring with excellent data tests.
The tests that most commonly score zero in practice:
- Data 3: no feature has a cost exceeding its benefit.
- Data 5: the system respects meta-level requirements (privacy, jurisdiction, retention).
- Model 6: a simpler model is not better.
- Infra 5: the model is tested via its serving API, not only in the training harness.
- Monitor 4: the model is not stale.
Monitor 4 is the axis where the paper reports the widest gap between perceived and actual readiness.
Behavioural testing
Ribeiro et al. (2020), Beyond Accuracy: Behavioral Testing of NLP Models with CheckList, argues that held-out accuracy is a single aggregate that hides systematic failure modes, and proposes testing capabilities rather than examples. Three test types:
- MFT (minimum functionality) — small targeted sets probing one capability.
- INV (invariance) — perturbations that must not change the prediction.
- DIR (directional expectation) — perturbations with a known expected direction of change.
Their user study found that experts testing a commercial sentiment model with CheckList discovered failure categories that had survived extensive standard evaluation. The framework is domain-general; the tabular translation is the table in the Developer tab.
Metamorphic testing
The underlying idea predates ML. Chen et al. (1998) introduced metamorphic relations for programs with no available oracle: rather than asserting an output, assert a relation between the outputs of related inputs.
For ML this dissolves the oracle problem. You may not know the correct probability for an applicant, but you know that raising income should not lower approval. The relation is checkable without ground truth, which is exactly the constraint ML testing operates under.
Useful relations for tabular models: monotonicity in a known-signed feature, permutation invariance over row order, scale invariance for tree models, and additivity checks for models with linear components.
Data validation as a first-class stage
TensorFlow Data Validation (Breck et al., 2019, Data Validation for Machine Learning) formalises three checks against a versioned schema:
- Single-batch validation — does today's data satisfy the declared schema (types, domains, required-ness, expected value ranges)?
- Inter-batch validation — does today's distribution differ materially from yesterday's, per feature?
- Training-serving skew detection — do the training and serving feature distributions agree?
The design decision worth copying is that the schema is a checked-in artefact with its own version history, not something inferred fresh each run. Inferring it fresh means the schema tracks whatever broke, which is the same failure mode as regenerating a golden set on failure.
Great Expectations and Pandera provide equivalents outside the TFX ecosystem.
Continuous training, and when it is wrong
CT — automatically retraining and redeploying on a trigger — is often presented as the mature end state. It is a genuine risk multiplier, and it should be gated on:
- A drift or decay signal validated against real degradation, not a p-value alone.
- Labels whose quality is verified, and which are not themselves generated by the current model.
- An offline evaluation gate strictly stronger than "the new model beat the old one on one metric".
- A canary or shadow stage before full traffic.
- Automatic rollback with a defined trigger and a defined observation window.
Absent any of these, scheduled manual retraining is the more defensible design. Automating a decision you cannot yet make correctly automates the error.
The offline-online gap
The metric that gates your pipeline is not the metric the business cares about, and the two correlate less than teams assume. Documented reasons:
- Selection bias in logged data — you only observe outcomes for actions the previous policy took.
- Position and presentation effects — in ranking, the interface shapes the outcome independent of relevance.
- Delayed and partial feedback — conversions attributed after the measurement window are invisible offline.
- Novelty and primacy effects — early online effects that decay, and which short A/B tests systematically misread.
Counterfactual evaluation (inverse propensity scoring, doubly robust estimators; Dudík et al., 2011) narrows the gap when the logging policy's action probabilities were recorded. If they were not recorded, they cannot be recovered afterwards — log propensities from day one or give up the option permanently.
Papers
- Breck et al., The ML Test Score, IEEE Big Data 2017 — research.google/pubs/pub46555
- Breck et al., Data Validation for Machine Learning, SysML 2019 — mlsys.org/Conferences/2019/doc/2019/167.pdf
- Ribeiro et al., Beyond Accuracy: Behavioral Testing of NLP Models with CheckList, ACL 2020 — arxiv.org/abs/2005.04118
- Chen, Cheung and Yiu, Metamorphic Testing: A New Approach for Generating Next Test Cases, 1998
- Dudík, Langford and Li, Doubly Robust Policy Evaluation and Learning, ICML 2011 — arxiv.org/abs/1103.4601
- Mitchell et al., Model Cards for Model Reporting, FAT* 2019 — arxiv.org/abs/1810.03993
- Gebru et al., Datasheets for Datasets, CACM 2021 — arxiv.org/abs/1803.09010
What to learn next
- Monitoring and model drift — the signal that should trigger a retrain in the first place.
- Docker for ML — the image this pipeline builds when the tests go green.
- Model deployment — the same discipline applied to language models.