Evaluation gates in CI
An evaluation gate turns many separate scores into one automatic ship-or-block decision, the same way a single cut-off mark decides a pass or fail regardless of which questions were answered well.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
An evaluation gate turns several separate scores into one automatic decision: ship, or block.
The analogy you have already lived
You have sat an exam with a pass mark. It does not matter how impressively you did on question one if your total falls under the cut-off — you still fail. One number, agreed in advance, decides the outcome. Nobody argues with the examiner about it on the day.
An evaluation gate is that same pass mark, applied to a model. Accuracy, a behavioural check, and a speed limit all get measured. A single rule, agreed on in advance, decides whether the model is allowed to ship.
Why it exists
Earlier lessons in this section built several different kinds of check: unit tests, behavioural tests, golden regressions, integration tests. Each one produces its own pass or fail.
Left as separate reports, it becomes a judgement call every time: three checks passed, one is borderline, is that good enough to ship today? Different people will answer that differently, on different days, under different amounts of pressure to ship quickly.
A gate removes the judgement call. The rule is written down once, in code, before anyone is under pressure. On release day, the computer applies it exactly the same way it did yesterday.
How it works
run every check compare each result
------------------ to an agreed threshold
accuracy: 0.91 accuracy >= 0.85 ? PASS
behavioural pass rate: 1.0 behaviour >= 1.00 ? PASS
latency p95: 143 ms latency <= 200 ? PASS
model size: 62 MB size <= 50 ? fails, but only a WARNING
|
v
every HARD gate passed?
/ \
yes no
| |
SHIP IT BLOCKEDA real example you have seen
A driving test has hard requirements — hit a pedestrian in the simulation, you fail, no matter how smoothly you parked. It also has softer notes — "a little hesitant at the junction" — that get written down without failing you outright. A good evaluation gate makes exactly that distinction on purpose.
Remember this
- A gate combines many separate checks into one ship-or-block decision, agreed on before release day, not during it.
- Hard gates block a release. Soft gates warn, but do not block — used for things worth watching, not worth stopping for.
- The threshold is a choice, made by people, in advance — not something the computer invents on its own.
What to learn next
- GitHub Actions for ML projects — wiring this gate script into a real pipeline that runs on every push.
- Is this improvement real? — deciding whether a metric change is signal or noise, before it ever reaches a threshold.
- Canary releases for models — applying the same ship-or-block logic to live traffic instead of an offline report.
Developer — Code and libraries.
Setup
pip install pytestNo external libraries needed for the gate itself — it reads a small JSON report and applies rules written directly in Python, which keeps the logic easy to read and easy to test on its own.
The gate script
"""Turns an evaluation report into one pass/fail decision for CI.
Run as: python gate.py eval_report.json
Exits 0 if every hard gate passes, 1 otherwise -- the exit code is what
a CI system actually reads to decide whether to continue the pipeline.
"""
import json
import sys
# Each gate: (report key, comparison, threshold, hard-or-soft)
GATES = [
("accuracy", "gte", 0.85, "hard"),
("behavioural_pass_rate", "gte", 1.00, "hard"),
("latency_p95_ms", "lte", 200, "hard"),
("model_size_mb", "lte", 50, "soft"),
]
def check(value, comparison, threshold):
return value >= threshold if comparison == "gte" else value <= threshold
def run_gate(report: dict) -> bool:
all_hard_passed = True
for key, comparison, threshold, kind in GATES:
value = report[key]
passed = check(value, comparison, threshold)
symbol = "PASS" if passed else "FAIL"
op = ">=" if comparison == "gte" else "<="
print(f"[{kind:4}] {symbol} {key:24} {value:>8} (needs {op} {threshold})")
if kind == "hard" and not passed:
all_hard_passed = False
return all_hard_passed
if __name__ == "__main__":
with open(sys.argv[1]) as f:
report = json.load(f)
ok = run_gate(report)
print()
print("RESULT:", "SHIP IT" if ok else "BLOCKED")
sys.exit(0 if ok else 1)Two real reports, run through the gate
{"accuracy": 0.91, "behavioural_pass_rate": 1.0, "latency_p95_ms": 143, "model_size_mb": 62}python gate.py eval_report_good.json[hard] PASS accuracy 0.91 (needs >= 0.85) [hard] PASS behavioural_pass_rate 1.0 (needs >= 1.0) [hard] PASS latency_p95_ms 143 (needs <= 200) [soft] FAIL model_size_mb 62 (needs <= 50) RESULT: SHIP IT
Notice the model is 62 MB against a 50 MB target — a soft gate, so it prints as FAIL but does not block the ship decision. Worth watching, not worth stopping the release for.
{"accuracy": 0.79, "behavioural_pass_rate": 0.83, "latency_p95_ms": 143, "model_size_mb": 30}python gate.py eval_report_bad.json[hard] FAIL accuracy 0.79 (needs >= 0.85) [hard] FAIL behavioural_pass_rate 0.83 (needs >= 1.0) [hard] PASS latency_p95_ms 143 (needs <= 200) [soft] PASS model_size_mb 30 (needs <= 50) RESULT: BLOCKED
Two hard gates fail, and the exit code (checked with echo $? right after running the script) is 1. That single exit code is the entire interface a CI system needs — a script that returns non-zero on failure is exactly what stops a pipeline.
Testing the gate logic itself
The gate is a piece of decision-making code, and it deserves the same kind of test as anything else in this section.
from gate import run_gate
def test_a_report_that_clears_every_hard_gate_passes():
report = {"accuracy": 0.91, "behavioural_pass_rate": 1.0, "latency_p95_ms": 143, "model_size_mb": 62}
assert run_gate(report) is True
def test_low_accuracy_alone_blocks_the_release():
report = {"accuracy": 0.79, "behavioural_pass_rate": 1.0, "latency_p95_ms": 143, "model_size_mb": 30}
assert run_gate(report) is False
def test_a_soft_gate_failing_does_not_block_the_release():
report = {"accuracy": 0.91, "behavioural_pass_rate": 1.0, "latency_p95_ms": 143, "model_size_mb": 999}
assert run_gate(report) is Truepytest test_gate.py -q... [100%] 3 passed in 0.01s
Testing gate logic on its own, separately from any real model, is worth doing precisely because the gate is the thing standing between a broken model and your users — it deserves to be at least as trustworthy as the checks that feed it.
Common mistakes
Setting the threshold at today's exact score. If accuracy is 0.910 today and the gate requires >= 0.910, a harmless retrain that lands at 0.909 blocks the release for no real reason. Set the line where you would genuinely refuse to ship, with a margin — this lesson's example uses 0.85 against a model scoring 0.91.
Making everything a hard gate. A gate that blocks on every metric, however minor, trains people to bypass it under deadline pressure. Reserve hard gates for the things that would be genuinely unsafe to ship; use soft gates for things worth tracking.
Making nothing a hard gate. The opposite failure is every bit as real — a gate that only ever warns has quietly become a suggestion, and a genuinely broken model will ship past it eventually.
No way to see why a release was blocked. The printed table above exists for exactly this reason — a bare exit code of 1 tells you nothing about which check failed. Always print the specific numbers and thresholds, not only the final verdict.
Combining scores into one blended number before gating. A single weighted score can hide a catastrophic failure in one metric behind a good score in another. Gate on each metric separately, as above, and only combine them for the final SHIP or BLOCKED headline.
Try it yourself
Add a fifth gate, hallucination_rate, with a hard threshold of <= 0.02. Write a test proving that a report with every other metric perfect still gets BLOCKED if this one gate fails — the property that makes a hard gate actually hard.
What to learn next
- GitHub Actions for ML projects — wiring this gate script into a real pipeline that runs on every push.
- Is this improvement real? — deciding whether a metric change is signal or noise, before it ever reaches a threshold.
- Canary releases for models — applying the same ship-or-block logic to live traffic instead of an offline report.
Researcher — Mathematics and papers.
Gates as a decision rule under multiple objectives
Formally, a gate is a function from a metric vector $m \in \mathbb{R}^k$ to a binary decision $d \in {\text{ship}, \text{block}}$, defined as a conjunction of per-metric thresholds over the hard-gate subset $H \subseteq {1, \ldots, k}$:
$$d = \text{ship} \iff \bigwedge_{i \in H} \left( m_i \geq \tau_i \text{ or } m_i \leq \tau_i \right)$$
This is deliberately a conjunction, not a weighted sum. A weighted-sum rule $\sum_i w_i m_i \geq \tau$ permits a large improvement in one metric to compensate for a catastrophic regression in another — precisely the failure mode Breck et al.'s (2017) ML Test Score avoids by defining its final score as the minimum across axes rather than an average, for the same reason argued in the developer section above.
Threshold placement and statistical noise
A hard threshold set exactly at the current model's score treats a single point estimate as ground truth, ignoring the sampling variance of the evaluation metric itself. If accuracy is measured on a held-out set of size $n$, its standard error is approximately $\sqrt{\hat{p}(1-\hat{p})/n}$; a threshold placed within one standard error of today's score will trigger false blocks from evaluation-set noise alone, indistinguishable from a real regression. The margin recommended in the developer section is a practical response to this: set the threshold outside the noise band, informed by statistical significance in ML rather than by the single number in front of you.
Multi-armed and sequential framings
Continuous deployment with an eval gate at each step is structurally a stopping rule in sequential decision theory: observe a metric, decide whether to proceed or halt, repeat. This connects the eval-gate pattern directly to the tooling in peeking and sequential testing once the gate is evaluated on live traffic (a canary) rather than an offline set — the offline gate in this lesson is the simpler, non-sequential special case.
Papers
- Breck et al., The ML Test Score: A Rubric for ML Production Readiness and Technical Debt Reduction, IEEE Big Data 2017 — the minimum-across-axes scoring rule this lesson's conjunction directly implements.
- Sculley et al., Hidden Technical Debt in Machine Learning Systems, NeurIPS 2015 — names "undeclared consumers" and threshold-free pipelines as sources of the exact ungoverned-release risk a gate is built to close.
What to learn next
- GitHub Actions for ML projects — wiring this gate script into a real pipeline that runs on every push.
- Is this improvement real? — deciding whether a metric change is signal or noise, before it ever reaches a threshold.
- Canary releases for models — applying the same ship-or-block logic to live traffic instead of an offline report.