Testing ML Code and CI

Evaluation gates in CI

An evaluation gate turns many separate scores into one automatic ship-or-block decision, the same way a single cut-off mark decides a pass or fail regardless of which questions were answered well.

On this page 7
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

An evaluation gate turns several separate scores into one automatic decision: ship, or block.

The analogy you have already lived

You have sat an exam with a pass mark. It does not matter how impressively you did on question one if your total falls under the cut-off — you still fail. One number, agreed in advance, decides the outcome. Nobody argues with the examiner about it on the day.

An evaluation gate is that same pass mark, applied to a model. Accuracy, a behavioural check, and a speed limit all get measured. A single rule, agreed on in advance, decides whether the model is allowed to ship.

Why it exists

Earlier lessons in this section built several different kinds of check: unit tests, behavioural tests, golden regressions, integration tests. Each one produces its own pass or fail.

Left as separate reports, it becomes a judgement call every time: three checks passed, one is borderline, is that good enough to ship today? Different people will answer that differently, on different days, under different amounts of pressure to ship quickly.

A gate removes the judgement call. The rule is written down once, in code, before anyone is under pressure. On release day, the computer applies it exactly the same way it did yesterday.

How it works

   run every check                     compare each result
   ------------------                   to an agreed threshold
   accuracy: 0.91                       accuracy   >= 0.85 ?   PASS
   behavioural pass rate: 1.0           behaviour  >= 1.00 ?   PASS
   latency p95: 143 ms                  latency    <= 200 ?   PASS
   model size: 62 MB                    size       <= 50  ?   fails, but only a WARNING
                                                |
                                                v
                                    every HARD gate passed?
                                          /            \
                                       yes              no
                                        |                |
                                    SHIP IT           BLOCKED

A real example you have seen

A driving test has hard requirements — hit a pedestrian in the simulation, you fail, no matter how smoothly you parked. It also has softer notes — "a little hesitant at the junction" — that get written down without failing you outright. A good evaluation gate makes exactly that distinction on purpose.

Remember this

  • A gate combines many separate checks into one ship-or-block decision, agreed on before release day, not during it.
  • Hard gates block a release. Soft gates warn, but do not block — used for things worth watching, not worth stopping for.
  • The threshold is a choice, made by people, in advance — not something the computer invents on its own.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install pytest

No external libraries needed for the gate itself — it reads a small JSON report and applies rules written directly in Python, which keeps the logic easy to read and easy to test on its own.

The gate script

gate.py
"""Turns an evaluation report into one pass/fail decision for CI.

Run as: python gate.py eval_report.json
Exits 0 if every hard gate passes, 1 otherwise -- the exit code is what
a CI system actually reads to decide whether to continue the pipeline.
"""
import json
import sys

# Each gate: (report key, comparison, threshold, hard-or-soft)
GATES = [
    ("accuracy",              "gte", 0.85, "hard"),
    ("behavioural_pass_rate", "gte", 1.00, "hard"),
    ("latency_p95_ms",        "lte", 200,  "hard"),
    ("model_size_mb",         "lte", 50,   "soft"),
]


def check(value, comparison, threshold):
    return value >= threshold if comparison == "gte" else value <= threshold


def run_gate(report: dict) -> bool:
    all_hard_passed = True
    for key, comparison, threshold, kind in GATES:
        value = report[key]
        passed = check(value, comparison, threshold)
        symbol = "PASS" if passed else "FAIL"
        op = ">=" if comparison == "gte" else "<="
        print(f"[{kind:4}] {symbol}  {key:24} {value:>8}  (needs {op} {threshold})")
        if kind == "hard" and not passed:
            all_hard_passed = False
    return all_hard_passed


if __name__ == "__main__":
    with open(sys.argv[1]) as f:
        report = json.load(f)
    ok = run_gate(report)
    print()
    print("RESULT:", "SHIP IT" if ok else "BLOCKED")
    sys.exit(0 if ok else 1)

Two real reports, run through the gate

eval_report_good.json
{"accuracy": 0.91, "behavioural_pass_rate": 1.0, "latency_p95_ms": 143, "model_size_mb": 62}
bash
python gate.py eval_report_good.json
Output
[hard] PASS  accuracy                     0.91  (needs >= 0.85)
[hard] PASS  behavioural_pass_rate         1.0  (needs >= 1.0)
[hard] PASS  latency_p95_ms                143  (needs <= 200)
[soft] FAIL  model_size_mb                  62  (needs <= 50)

RESULT: SHIP IT

Notice the model is 62 MB against a 50 MB target — a soft gate, so it prints as FAIL but does not block the ship decision. Worth watching, not worth stopping the release for.

eval_report_bad.json
{"accuracy": 0.79, "behavioural_pass_rate": 0.83, "latency_p95_ms": 143, "model_size_mb": 30}
bash
python gate.py eval_report_bad.json
Output
[hard] FAIL  accuracy                     0.79  (needs >= 0.85)
[hard] FAIL  behavioural_pass_rate        0.83  (needs >= 1.0)
[hard] PASS  latency_p95_ms                143  (needs <= 200)
[soft] PASS  model_size_mb                  30  (needs <= 50)

RESULT: BLOCKED

Two hard gates fail, and the exit code (checked with echo $? right after running the script) is 1. That single exit code is the entire interface a CI system needs — a script that returns non-zero on failure is exactly what stops a pipeline.

Testing the gate logic itself

The gate is a piece of decision-making code, and it deserves the same kind of test as anything else in this section.

test_gate.py
from gate import run_gate


def test_a_report_that_clears_every_hard_gate_passes():
    report = {"accuracy": 0.91, "behavioural_pass_rate": 1.0, "latency_p95_ms": 143, "model_size_mb": 62}
    assert run_gate(report) is True


def test_low_accuracy_alone_blocks_the_release():
    report = {"accuracy": 0.79, "behavioural_pass_rate": 1.0, "latency_p95_ms": 143, "model_size_mb": 30}
    assert run_gate(report) is False


def test_a_soft_gate_failing_does_not_block_the_release():
    report = {"accuracy": 0.91, "behavioural_pass_rate": 1.0, "latency_p95_ms": 143, "model_size_mb": 999}
    assert run_gate(report) is True
bash
pytest test_gate.py -q
Output
...                                                                      [100%]
3 passed in 0.01s

Testing gate logic on its own, separately from any real model, is worth doing precisely because the gate is the thing standing between a broken model and your users — it deserves to be at least as trustworthy as the checks that feed it.

Common mistakes

Setting the threshold at today's exact score. If accuracy is 0.910 today and the gate requires >= 0.910, a harmless retrain that lands at 0.909 blocks the release for no real reason. Set the line where you would genuinely refuse to ship, with a margin — this lesson's example uses 0.85 against a model scoring 0.91.

Making everything a hard gate. A gate that blocks on every metric, however minor, trains people to bypass it under deadline pressure. Reserve hard gates for the things that would be genuinely unsafe to ship; use soft gates for things worth tracking.

Making nothing a hard gate. The opposite failure is every bit as real — a gate that only ever warns has quietly become a suggestion, and a genuinely broken model will ship past it eventually.

No way to see why a release was blocked. The printed table above exists for exactly this reason — a bare exit code of 1 tells you nothing about which check failed. Always print the specific numbers and thresholds, not only the final verdict.

Combining scores into one blended number before gating. A single weighted score can hide a catastrophic failure in one metric behind a good score in another. Gate on each metric separately, as above, and only combine them for the final SHIP or BLOCKED headline.

Try it yourself

Add a fifth gate, hallucination_rate, with a hard threshold of <= 0.02. Write a test proving that a report with every other metric perfect still gets BLOCKED if this one gate fails — the property that makes a hard gate actually hard.

What to learn next

Researcher — Mathematics and papers.

Gates as a decision rule under multiple objectives

Formally, a gate is a function from a metric vector $m \in \mathbb{R}^k$ to a binary decision $d \in {\text{ship}, \text{block}}$, defined as a conjunction of per-metric thresholds over the hard-gate subset $H \subseteq {1, \ldots, k}$:

$$d = \text{ship} \iff \bigwedge_{i \in H} \left( m_i \geq \tau_i \text{ or } m_i \leq \tau_i \right)$$

This is deliberately a conjunction, not a weighted sum. A weighted-sum rule $\sum_i w_i m_i \geq \tau$ permits a large improvement in one metric to compensate for a catastrophic regression in another — precisely the failure mode Breck et al.'s (2017) ML Test Score avoids by defining its final score as the minimum across axes rather than an average, for the same reason argued in the developer section above.

Threshold placement and statistical noise

A hard threshold set exactly at the current model's score treats a single point estimate as ground truth, ignoring the sampling variance of the evaluation metric itself. If accuracy is measured on a held-out set of size $n$, its standard error is approximately $\sqrt{\hat{p}(1-\hat{p})/n}$; a threshold placed within one standard error of today's score will trigger false blocks from evaluation-set noise alone, indistinguishable from a real regression. The margin recommended in the developer section is a practical response to this: set the threshold outside the noise band, informed by statistical significance in ML rather than by the single number in front of you.

Multi-armed and sequential framings

Continuous deployment with an eval gate at each step is structurally a stopping rule in sequential decision theory: observe a metric, decide whether to proceed or halt, repeat. This connects the eval-gate pattern directly to the tooling in peeking and sequential testing once the gate is evaluated on live traffic (a canary) rather than an offline set — the offline gate in this lesson is the simpler, non-sequential special case.

Papers

  • Breck et al., The ML Test Score: A Rubric for ML Production Readiness and Technical Debt Reduction, IEEE Big Data 2017 — the minimum-across-axes scoring rule this lesson's conjunction directly implements.
  • Sculley et al., Hidden Technical Debt in Machine Learning Systems, NeurIPS 2015 — names "undeclared consumers" and threshold-free pipelines as sources of the exact ungoverned-release risk a gate is built to close.

What to learn next

Related terms

What to learn next

These follow on from what you just read.

  • Testing ML Code and CI

    GitHub Actions for ML projects

    GitHub Actions is the machine that actually runs your checks on every push — and an ML pipeline needs a few habits ordinary web-app CI does not, like skipping an expensive retrain when only the docs changed.

  • Testing ML Code and CI

    Automated retraining pipelines

    An automated retraining pipeline checks for enough new data, retrains, and only replaces the live model if it passes the same gate every other release has to clear — the same discipline as testing a fresh batch of curd before trusting it as tomorrow's starter.

  • Releasing Models Safely

    Shadow deployment

    A shadow deployment lets a new model answer every real request in secret, alongside the model actually serving users, so you can compare them honestly without ever risking a wrong answer reaching anyone.