Testing ML Code and CI

Automated retraining pipelines

An automated retraining pipeline checks for enough new data, retrains, and only replaces the live model if it passes the same gate every other release has to clear — the same discipline as testing a fresh batch of curd before trusting it as tomorrow's starter.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

An automated retraining pipeline periodically trains a fresh model on new data, and only keeps it if it passes a quality check — never automatically, unconditionally.

The analogy you have already lived

You have set curd at home using a spoonful of yesterday's curd as a starter. You do this again and again, but you always check the new batch actually set properly before you rely on it as tomorrow's starter. If a batch does not set — the milk was too cold, or the starter was weak — you do not serve it. You do not use it to start the next batch either. You go back to what you know is good.

Automated retraining is that same routine, run by a computer. New data arrives, a fresh model gets trained from it — but it only becomes the model everyone uses if it actually passes the check.

Why it exists

A model trained once, on data from six months ago, slowly falls out of step with the world. New patterns appear. Old ones fade. Monitoring and model drift is how you notice this is happening.

Retraining by hand, every time, does not scale. It depends on someone remembering to do it, at the right time, correctly, under whatever pressure that week happens to bring.

An automated pipeline removes the remembering. It checks on a schedule, and retrains when there is enough new data to be worth it. It never lets a bad batch — the model equivalent of curd that did not set — replace the one currently serving real users.

How it works

   on a schedule (say, every night)
              |
              v
   is there enough NEW labelled data since last time?
        /                              \
      no                               yes
       |                                 |
   SKIP, try again                  train a fresh model
   next scheduled run                     |
                                            v
                          test it against the same frozen exam paper
                          every release has to pass  (see: evaluation gates)
                                    /                    \
                                 fails                  passes
                                   |                       |
                          BLOCKED: keep the          PROMOTED: this becomes
                          old model running          the new live model

A real example you have seen

A spam filter that still catches new kinds of spam months after launch, without anyone visibly updating it, is running exactly this loop in the background. New labelled examples — mail people marked as spam — come in, and a model retrains. Only a model that still catches spam without flagging real mail gets to replace the one currently running.

The honest part

Automating retraining without automating the check is one of the more dangerous mistakes in this entire area. It turns a quiet data problem into an automatic, unattended way to break production. The gate is not optional decoration on this pipeline. It is the entire safety mechanism.

Remember this

  • Retraining runs on a schedule or a trigger, not only when a person remembers to.
  • A new model is only promoted after it passes the same quality gate as any other release.
  • If a batch fails the check, the pipeline keeps the previous model running — automation should never mean "ship it no matter what."

What to learn next

  • Canary releases for models — sending a newly promoted model to a small slice of real traffic before it takes over completely.
  • Measuring data drift — a smarter trigger for retraining than a fixed schedule.
  • Model registries — a more durable, auditable home for every model this pipeline promotes than a single overwritten file.

Developer — Code and libraries.

Setup

bash
pip install scikit-learn pandas joblib

sqlite3 is part of the Python standard library — no separate install needed.

The pipeline

A small, complete version of the loop above: check for new rows in a database, retrain if there are enough, evaluate against the same frozen golden set from CI/CD for machine learning, and only promote if it passes.

pipeline.py
"""A small, real automated-retraining pipeline. Meant to be run on a schedule."""
import shutil
import sqlite3
import sys

import joblib
import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

MIN_NEW_ROWS = 20
MIN_ACCURACY = 0.85
COLUMNS = ["income", "years", "age"]

# The same frozen 20-row exam paper from ci-cd-for-ml.md -- never regenerated.
GOLDEN = pd.DataFrame([
    (77.5, 4.4, 44, 1), (46.0, 9.5, 27, 1), (78.0, 7.9, 60, 1),
    (58.6, 8.7, 63, 1), (57.3, 1.7, 28, 0), (21.2, 0.7, 30, 0),
    (78.2, 6.0, 39, 1), (5.5, 1.7, 42, 0), (24.0, 7.3, 51, 1),
    (37.6, 4.1, 36, 1), (63.5, 5.3, 57, 1), (19.8, 9.4, 48, 1),
    (69.7, 5.2, 64, 1), (78.8, 1.1, 29, 1), (17.3, 1.6, 58, 0),
    (49.8, 5.5, 29, 1), (5.7, 5.2, 60, 0), (34.0, 6.4, 53, 1),
    (8.3, 4.0, 51, 0), (76.7, 6.5, 25, 1),
], columns=COLUMNS + ["repaid"])


def run(db_path: str = "applicants.db") -> str:
    con = sqlite3.connect(db_path)
    new_rows = con.execute(
        "SELECT COUNT(*) FROM applicants WHERE used_in_training = 0"
    ).fetchone()[0]
    print(f"new labelled rows since last training: {new_rows}")

    if new_rows < MIN_NEW_ROWS:
        print(f"SKIPPED: fewer than {MIN_NEW_ROWS} new rows, not worth a retrain yet")
        con.close()
        return "skipped"

    all_data = pd.read_sql("SELECT * FROM applicants", con)
    model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000))
    model.fit(all_data[COLUMNS], all_data["repaid"])

    accuracy = (model.predict(GOLDEN[COLUMNS]) == GOLDEN["repaid"]).mean()
    print(f"trained on {len(all_data)} rows, golden-set accuracy: {accuracy:.3f}")

    if accuracy < MIN_ACCURACY:
        print(f"BLOCKED: accuracy {accuracy:.3f} is below the {MIN_ACCURACY} gate")
        print("keeping the previous model.joblib; new rows stay unmarked for next run")
        con.close()
        return "blocked"

    joblib.dump(model, "model.joblib.new")
    shutil.move("model.joblib.new", "model.joblib")
    con.execute("UPDATE applicants SET used_in_training = 1 WHERE used_in_training = 0")
    con.commit()
    con.close()
    print("PROMOTED: model.joblib updated, new rows marked as used")
    return "promoted"


if __name__ == "__main__":
    result = run()
    sys.exit(0 if result != "blocked" else 1)

Three real runs, three real outcomes

Every run below reads from applicants.db, an ordinary SQLite file with one applicants table (income, years, age, repaid, used_in_training). This seeds it with 60 clean, unused rows:

seed_clean_batch.py
"""Run once by hand to seed applicants.db before the first two runs below."""
import sqlite3

import numpy as np
import pandas as pd


def make_batch(n: int, seed: int) -> pd.DataFrame:
    rng = np.random.RandomState(seed)
    X = pd.DataFrame({
        "income": rng.uniform(5, 80, n).round(1),
        "years": rng.uniform(0, 10, n).round(1),
        "age": rng.randint(21, 65, n).astype(float),
    })
    score = 0.05 * X["income"] + 0.35 * X["years"] + 0.01 * X["age"] - 3.5
    X["repaid"] = (score + rng.normal(0, 0.8, n) > 0).astype(int)
    X["used_in_training"] = 0
    return X


con = sqlite3.connect("applicants.db")
make_batch(60, seed=0).to_sql("applicants", con, if_exists="replace", index=False)
con.close()

With a threshold of 20 new rows:

bash
python seed_clean_batch.py
python pipeline.py
Output
new labelled rows since last training: 60
trained on 60 rows, golden-set accuracy: 0.950
PROMOTED: model.joblib updated, new rows marked as used

Run again immediately, with no new data arriving in between:

bash
python pipeline.py
Output
new labelled rows since last training: 0
SKIPPED: fewer than 20 new rows, not worth a retrain yet

Now a separate, genuinely broken scenario: a fresh applicants.db, reseeded with 40 rows where the label column is pure noise, unrelated to any feature — a realistic stand-in for a broken upstream labelling job:

seed_broken_batch.py
"""Run once by hand to reseed applicants.db with a broken labelling job."""
import sqlite3

import numpy as np
import pandas as pd


def make_broken_batch(n: int, seed: int) -> pd.DataFrame:
    rng = np.random.RandomState(seed)
    X = pd.DataFrame({
        "income": rng.uniform(5, 80, n).round(1),
        "years": rng.uniform(0, 10, n).round(1),
        "age": rng.randint(21, 65, n).astype(float),
    })
    X["repaid"] = rng.randint(0, 2, n)  # unrelated to income/years/age -- the bug
    X["used_in_training"] = 0
    return X


con = sqlite3.connect("applicants.db")
make_broken_batch(40, seed=7).to_sql("applicants", con, if_exists="replace", index=False)
con.close()
bash
python seed_broken_batch.py
python pipeline.py
Output
new labelled rows since last training: 40
trained on 40 rows, golden-set accuracy: 0.450
BLOCKED: accuracy 0.450 is below the 0.85 gate
keeping the previous model.joblib; new rows stay unmarked for next run

Run again right after: the same 40 rows are still unused, and the pipeline reports the same block, because a blocked run never marks its input data as consumed. That single design choice is what makes this pipeline safe to leave unattended — a bad batch does not quietly disappear, it sits there until someone looks at why the accuracy dropped.

Line-by-line: why the gate reuses the same golden set

The GOLDEN set here is the exact same twenty frozen examples used in CI/CD for machine learning. This is deliberate, not laziness — a retraining pipeline and a human-triggered release should be held to the identical bar. If they used different check sets, a model could pass whichever gate happened to be weaker, and the "automated" path would quietly become the easier way to ship a worse model.

Common mistakes

Promoting on any new data, however little. Retraining on five new rows adds noise more than signal, and rebuilding a model on a tiny increment is usually not worth the risk. MIN_NEW_ROWS exists specifically to wait for a batch large enough to matter.

No memory of what has already been used. Without the used_in_training flag, every run would retrain on the same data as last time plus a little more, wasting compute without learning anything new — or worse, a blocked batch would vanish from consideration on the next run instead of being retried.

Marking data as used even when the run is blocked. This is the single most important line in the whole script: the UPDATE that marks rows as used only runs in the promoted path. A pipeline that marked data as consumed regardless of outcome would lose a bad batch's evidence exactly when you need it most, for a postmortem.

No human ever looking at a blocked run. An unattended pipeline that silently blocks forever, with nobody told, is not much better than one that silently ships. Wire the blocked branch to a real alert — see alerts people do not ignore — not a log line nobody reads.

Retraining on a schedule when the real trigger should be drift. A nightly schedule is simple and often good enough. A better trigger, when you can build it, is measuring data drift directly — retrain when the world has actually moved, not on a fixed clock that might fire too often or too rarely.

Try it yourself

Change MIN_ACCURACY to 0.99 and rerun the first, clean scenario. Even good, real data will now fail to clear the bar. This is worth doing once: it makes concrete that a gate set unrealistically high blocks everything, which is its own kind of failure — a pipeline that never promotes anything gives you none of the benefit of automating retraining in the first place.

What to learn next

  • Canary releases for models — sending a newly promoted model to a small slice of real traffic before it takes over completely.
  • Measuring data drift — a smarter trigger for retraining than a fixed schedule.
  • Model registries — a more durable, auditable home for every model this pipeline promotes than a single overwritten file.

Researcher — Mathematics and papers.

When continuous training is the wrong design

The researcher section of CI/CD for machine learning lists five preconditions for safe continuous training: a validated drift or decay signal, verified label quality, an offline gate strictly stronger than "beat the old model," a canary or shadow stage before full traffic, and automatic rollback with a defined trigger. This lesson's pipeline satisfies the third precondition directly (the golden-set gate) and explicitly does not attempt the fourth or fifth — the promoted model here becomes the new model.joblib immediately, with no staged rollout. Wiring this pipeline's output into canary releases for models is what closes that remaining gap, and is the appropriate next step before trusting this pattern with real user traffic.

Feedback loops and label quality

The blocked scenario above uses labels drawn independently of the features — a clean stand-in for a real and common failure: an upstream labelling process that silently breaks (a join key changes, an event stops firing) while continuing to produce syntactically valid rows. This class of failure is invisible to schema checks and only visible to a quality gate evaluated against ground truth the pipeline itself did not produce — precisely why the golden set here is a small, separately curated, human-verified artifact rather than a held-out slice of the same incoming stream. A gate validated against data influenced by the model's own prior predictions can enter a feedback loop where degrading quality reinforces itself; feedback loops in production covers this failure mode directly.

Idempotency and at-least-once scheduling

A production scheduler (cron, a workflow orchestrator, a serverless trigger) generally offers at-least-once execution, not exactly-once — a retry after a timeout can invoke the pipeline twice for what was meant to be a single run. The used_in_training flag makes this pipeline naturally idempotent under that condition: a second, redundant invocation immediately after a successful promotion sees zero new rows and safely no-ops, rather than retraining twice or double-counting data. Designing a scheduled job to be safe under duplicate invocation is a general distributed-systems property, not an ML-specific one, but it is easy to lose sight of in an ML pipeline where "it's only a quick retrain" can feel harmless when it is not free.

Papers

  • Sculley et al., Hidden Technical Debt in Machine Learning Systems, NeurIPS 2015 — names feedback loops and "undeclared consumers" of a pipeline's own output as central sources of the debt this design tries to avoid.
  • Breck et al., The ML Test Score, IEEE Big Data 2017 — Infra Test 2, "model quality is validated before serving," is exactly the property this pipeline's gate implements as a hard block, not an advisory check.

What to learn next

  • Canary releases for models — sending a newly promoted model to a small slice of real traffic before it takes over completely.
  • Measuring data drift — a smarter trigger for retraining than a fixed schedule.
  • Model registries — a more durable, auditable home for every model this pipeline promotes than a single overwritten file.

What to learn next

These follow on from what you just read.

  • Releasing Models Safely

    Shadow deployment

    A shadow deployment lets a new model answer every real request in secret, alongside the model actually serving users, so you can compare them honestly without ever risking a wrong answer reaching anyone.

  • Releasing Models Safely

    Canary releases for models

    A canary release gives a new model a small, real slice of traffic and watches it closely, the way miners once carried a caged canary underground — small warning, before anyone bigger got hurt.

  • Releasing Models Safely

    Blue-green model deployments

    A blue-green deploy keeps two complete, fully warmed-up environments and flips every request between them with one instant switch, the way a home changeover switch moves the whole house from grid power to a generator, never a blend of both.