MLOps

CI/CD for machine learning

A machine that rebuilds, retests and repackages your model on every change — and refuses to ship it when a check fails, no matter who is asking.

Read these first

On this page 10
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. What makes this harder than normal software
  5. How it works
  6. The frozen exam paper
  7. Where you have already benefited from this
  8. The honest part
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

CI/CD is a machine that rebuilds and rechecks everything on every change, and refuses to ship when a check fails.

CI is short for continuous integration — every change gets combined with everyone else's and tested, straight away. CD is continuous delivery — if the tests pass, the result is packaged and sent onward automatically.

The analogy you have already lived

You have walked through a security check at a mall or an airport. Your bag goes on the belt, through the scanner, every single time.

Nobody is exempt. Not the regular visitor, not the manager, not the person who is late. The machine does not get tired at 11 p.m. on a Friday and wave someone through.

That is CI. The value is not that it is clever. The value is that it cannot be talked out of it.

Why it exists

Picture the version without it.

Somebody makes a change on Friday evening. They mean to run the tests. They are in a hurry, and the tests take eleven minutes, and it is a small change.

On Monday the model returns nothing but zeros. Nobody knows which of the nine changes made that weekend caused it. Three people spend a day finding out.

CI removes the decision. Change goes in, checks run, result is green or red. There is nothing to remember and nothing to skip.

What makes this harder than normal software

Ordinary software testing asks one question: does the code do what I said?

Machine learning needs three more questions. Skipping them is why so many pipelines pass every test and still ship a broken model.

Is the data still what we agreed? The right columns, the right types, values in sensible ranges, not suddenly half empty.

Is the model still good enough? Score it on a frozen set of examples with known answers. If it drops below an agreed line, stop.

Does the model still behave sensibly? Some things must always be true. Raising an applicant's income must never lower their approval score. A rule like that catches disasters no accuracy number will.

How it works

   you push a change
          |
          v
   install the pinned libraries
          |
   train the model
          |
   test the CODE     does it run, does it return the right shape
   test the DATA     right columns, no floods of missing values
   test the MODEL    score on frozen examples, sanity rules hold
          |
    all green?  ---- no ---->  STOP. Nothing ships. Somebody gets told.
          |
         yes
          v
   build the sealed box  ->  deploy to a copy  ->  deploy for real

Everything above the "STOP" line is CI. Everything below it is CD.

Most teams get enormous value from the top half alone, and it is where you should start.

The frozen exam paper

The single most useful thing on this page is small enough to build today.

Take twenty real examples where you know the correct answer. Save them in a file. Check that file into version control and never regenerate it.

Every time the model changes, it sits that same exam. If the score falls below your line, the change does not ship.

It is twenty rows. It catches the wrong data file and the flipped label column. It catches a feature that stopped being computed. It catches a retrain that quietly failed halfway.

Where you have already benefited from this

Every app on your phone that updates without breaking. Every website that changes overnight and still works in the morning.

You never see CI/CD. You see its absence, on the days something ships broken.

The honest part

CI for machine learning is slower and more irritating than CI for ordinary code.

Training takes minutes or hours, not seconds. Model tests can be flaky, because training has randomness in it. A test that passes four times out of five is worse than no test. People learn to re-run it until it goes green.

The cures are real, and they take discipline. Fix your random seeds. Set your threshold with a margin below today's score, not at it. Keep the slow full retrain on a schedule, and run the fast checks on every change.

Remember this

  • CI runs the checks automatically, so nobody can skip them in a hurry.
  • Machine learning needs tests for data and model, not only code.
  • A frozen set of twenty known examples is the cheapest safety net you will ever build.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install pytest scikit-learn pandas joblib

This lesson continues from model serving. It uses the model.joblib that the train.py there produces. Run that script first, or the fixture below has nothing to load.

The test file

Six tests. Each one catches a failure that has actually shipped somewhere.

test_model.py
import joblib
import numpy as np
import pandas as pd
import pytest

COLUMNS = ["income", "years", "age"]
MIN_ACCURACY = 0.85

# A frozen set of 20 applicants with known outcomes. Checked into git, never
# regenerated. This is the exam paper the model has to keep passing.
GOLDEN = pd.DataFrame([
    (77.5, 4.4, 44, 1), (46.0, 9.5, 27, 1), (78.0, 7.9, 60, 1),
    (58.6, 8.7, 63, 1), (57.3, 1.7, 28, 0), (21.2, 0.7, 30, 0),
    (78.2, 6.0, 39, 1), (5.5, 1.7, 42, 0), (24.0, 7.3, 51, 1),
    (37.6, 4.1, 36, 1), (63.5, 5.3, 57, 1), (19.8, 9.4, 48, 1),
    (69.7, 5.2, 64, 1), (78.8, 1.1, 29, 1), (17.3, 1.6, 58, 0),
    (49.8, 5.5, 29, 1), (5.7, 5.2, 60, 0), (34.0, 6.4, 53, 1),
    (8.3, 4.0, 51, 0), (76.7, 6.5, 25, 1),
], columns=COLUMNS + ["repaid"])


@pytest.fixture(scope="module")
def model():
    return joblib.load("model.joblib")


def test_model_expects_exactly_the_agreed_columns(model):
    assert list(model.feature_names_in_) == COLUMNS


def test_every_score_is_a_usable_probability(model):
    p = model.predict_proba(GOLDEN[COLUMNS])[:, 1]
    assert np.isfinite(p).all(), "a NaN reached the caller"
    assert ((p >= 0.0) & (p <= 1.0)).all()


def test_more_income_never_lowers_the_score(model):
    """A directional expectation. Nothing else changes, only income rises."""
    ladder = pd.DataFrame({"income": np.arange(10.0, 80.0, 5.0),
                           "years": 4.0, "age": 35.0})
    p = model.predict_proba(ladder[COLUMNS])[:, 1]
    assert np.all(np.diff(p) >= -1e-9), "score fell when income rose"


def test_accuracy_on_the_frozen_set(model):
    accuracy = (model.predict(GOLDEN[COLUMNS]) == GOLDEN["repaid"]).mean()
    assert accuracy >= MIN_ACCURACY, f"accuracy fell to {accuracy:.3f}"


def test_a_very_strong_applicant_is_approved(model):
    strong = pd.DataFrame([{"income": 75.0, "years": 9.0, "age": 40.0}])
    assert model.predict(strong[COLUMNS])[0] == 1


def test_missing_values_are_refused_not_guessed(model):
    broken = pd.DataFrame([{"income": np.nan, "years": 4.0, "age": 35.0}])
    with pytest.raises(ValueError):
        model.predict(broken[COLUMNS])
bash
pytest test_model.py -q
Output
......                                                                   [100%]
6 passed in 0.79s

Now break something on purpose

This is the part worth doing rather than reading. Retrain with the label column inverted. That is a real and common bug: an upstream table changes 1 from meaning "repaid" to meaning "defaulted".

python
# add this line to train.py, replacing the fit call
model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000)).fit(X, 1 - y)
Output
..FFF.                                                                   [100%]
=========================== short test summary info ===========================
FAILED test_model.py::test_more_income_never_lowers_the_score - AssertionError: score fell when i...
FAILED test_model.py::test_accuracy_on_the_frozen_set - AssertionError: accuracy fell to 0.050
FAILED test_model.py::test_a_very_strong_applicant_is_approved - assert 0 == 1
3 failed, 3 passed in 0.83s

Line widths and the truncation with ... depend on your terminal, so those will differ. What matters is which tests moved.

Three different tests caught one bug, and each one describes it differently. Accuracy fell to 0.050. The score fell when income rose. An applicant with every positive signal was rejected. Any one of them alone would have stopped the release.

Notice which tests still passed. The schema was fine. The probabilities were valid numbers between zero and one. A suite that checked only shapes and types would have shown six green ticks. It would have shipped a model that is exactly backwards.

Wiring it to every push

.github/workflows/model-checks.yml
name: model checks

on:
  push:
  pull_request:

jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4

      - uses: actions/setup-python@v5
        with:
          python-version: "3.11"
          cache: pip

      # Pinned versions, so a library release cannot change the result
      # of a test run without somebody editing a file first.
      - name: Install dependencies
        run: pip install -r requirements.txt

      - name: Train
        run: python train.py

      - name: Test the model
        run: pytest -q

      # Only reached when every test above passed.
      - name: Build the image
        run: docker build -t loan-model:${{ github.sha }} .

No output block here. The log is a live web page with timings and collapsible steps, and reproducing it as text would be a fiction.

The important property is the ordering. docker build sits after pytest, and GitHub Actions stops a job at the first failing step. A red test means no image is ever built, so a broken model cannot reach a registry by accident.

The four kinds of ML test, named

Worth knowing the vocabulary, because the literature uses it.

KindQuestionExample above
SchemaRight columns, right types?test_model_expects_exactly_the_agreed_columns
Minimum functionalityDoes the easiest possible case work?test_a_very_strong_applicant_is_approved
Directional expectationDoes changing one input move the output the right way?test_more_income_never_lowers_the_score
InvarianceDoes a change that should not matter leave the output alone?not shown — add one for a field the model must ignore

Ribeiro et al. (2020) formalised these as behavioural testing, in the paper that introduced CheckList. The idea transfers cleanly from language models to tabular ones.

Common mistakes

Setting the threshold at today's exact score. If the model scores 0.87 and you assert >= 0.87, the next harmless retrain fails the build. Set the line where you would genuinely refuse to ship — here, 0.85 against a model scoring 0.95.

Regenerating the golden set when a test fails. This is the ML equivalent of deleting the failing test. If the frozen examples are wrong, fix them deliberately, in their own commit, with the reason written down.

Retraining from scratch on every commit. For anything larger than a toy this makes CI unusably slow. Split it: fast tests against the last committed model on every push, full retrain nightly or on demand.

Unpinned dependencies in CI. Without pins, a library released this morning can turn your build red with no change from you. Worse, it can turn it green when it should be red.

Testing only the model. The serving layer needs tests too. The TestClient checks from model serving run in under a second and catch the validation bugs that models never see.

No path to roll back. Automated deployment without a one-command rollback is automated risk. Keep the previous image tagged and know the command before you need it.

Try it yourself

Add an invariance test. Give the model two applicants whose only difference is a field you believe it should ignore, and assert the scores match.

If your model has no such field, that finding is more valuable than the test. It means every column you collect is influencing decisions, including any you would struggle to defend.

What to learn next

Researcher — Mathematics and papers.

The rubric

Breck et al. (2017), The ML Test Score, remains the most useful checklist in the field because it is specific enough to argue with. Twenty-eight tests over four axes, each scored 0, 0.5 or 1, with the final score being the minimum across axes — deliberately, so a team cannot compensate for absent monitoring with excellent data tests.

The tests that most commonly score zero in practice:

  • Data 3: no feature has a cost exceeding its benefit.
  • Data 5: the system respects meta-level requirements (privacy, jurisdiction, retention).
  • Model 6: a simpler model is not better.
  • Infra 5: the model is tested via its serving API, not only in the training harness.
  • Monitor 4: the model is not stale.

Monitor 4 is the axis where the paper reports the widest gap between perceived and actual readiness.

Behavioural testing

Ribeiro et al. (2020), Beyond Accuracy: Behavioral Testing of NLP Models with CheckList, argues that held-out accuracy is a single aggregate that hides systematic failure modes, and proposes testing capabilities rather than examples. Three test types:

  • MFT (minimum functionality) — small targeted sets probing one capability.
  • INV (invariance) — perturbations that must not change the prediction.
  • DIR (directional expectation) — perturbations with a known expected direction of change.

Their user study found that experts testing a commercial sentiment model with CheckList discovered failure categories that had survived extensive standard evaluation. The framework is domain-general; the tabular translation is the table in the Developer tab.

Metamorphic testing

The underlying idea predates ML. Chen et al. (1998) introduced metamorphic relations for programs with no available oracle: rather than asserting an output, assert a relation between the outputs of related inputs.

For ML this dissolves the oracle problem. You may not know the correct probability for an applicant, but you know that raising income should not lower approval. The relation is checkable without ground truth, which is exactly the constraint ML testing operates under.

Useful relations for tabular models: monotonicity in a known-signed feature, permutation invariance over row order, scale invariance for tree models, and additivity checks for models with linear components.

Data validation as a first-class stage

TensorFlow Data Validation (Breck et al., 2019, Data Validation for Machine Learning) formalises three checks against a versioned schema:

  • Single-batch validation — does today's data satisfy the declared schema (types, domains, required-ness, expected value ranges)?
  • Inter-batch validation — does today's distribution differ materially from yesterday's, per feature?
  • Training-serving skew detection — do the training and serving feature distributions agree?

The design decision worth copying is that the schema is a checked-in artefact with its own version history, not something inferred fresh each run. Inferring it fresh means the schema tracks whatever broke, which is the same failure mode as regenerating a golden set on failure.

Great Expectations and Pandera provide equivalents outside the TFX ecosystem.

Continuous training, and when it is wrong

CT — automatically retraining and redeploying on a trigger — is often presented as the mature end state. It is a genuine risk multiplier, and it should be gated on:

  1. A drift or decay signal validated against real degradation, not a p-value alone.
  2. Labels whose quality is verified, and which are not themselves generated by the current model.
  3. An offline evaluation gate strictly stronger than "the new model beat the old one on one metric".
  4. A canary or shadow stage before full traffic.
  5. Automatic rollback with a defined trigger and a defined observation window.

Absent any of these, scheduled manual retraining is the more defensible design. Automating a decision you cannot yet make correctly automates the error.

The offline-online gap

The metric that gates your pipeline is not the metric the business cares about, and the two correlate less than teams assume. Documented reasons:

  • Selection bias in logged data — you only observe outcomes for actions the previous policy took.
  • Position and presentation effects — in ranking, the interface shapes the outcome independent of relevance.
  • Delayed and partial feedback — conversions attributed after the measurement window are invisible offline.
  • Novelty and primacy effects — early online effects that decay, and which short A/B tests systematically misread.

Counterfactual evaluation (inverse propensity scoring, doubly robust estimators; Dudík et al., 2011) narrows the gap when the logging policy's action probabilities were recorded. If they were not recorded, they cannot be recovered afterwards — log propensities from day one or give up the option permanently.

Papers

What to learn next