Testing ML Code and CI

Testing output you cannot predict

When a model's exact wording changes every time, you test the shape and the rules the output must always obey — the way you judge two cups of chai by heat and sweetness, not by counting identical tea leaves.

On this page 7
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

When exact output cannot be predicted, you test the rules it must always obey, instead of one exact expected answer.

The analogy you have already lived

You have ordered chai from the same stall twice. The two cups were not identical — a little more or less sugar, poured slightly differently. You did not complain about that. You would complain if one arrived cold, or without any tea in it at all.

You were testing properties — hot, sweet, actually tea — not testing for one exact liquid. Large language models are the same. Ask the same question twice and the wording often differs, because these models pick from many likely next words rather than one fixed answer.

Why it exists

The golden output test from the previous lesson checks that an output matches a saved one, exactly. That works for a deterministic function — same input, same output, always.

A model that samples its next words does not give the same output twice, even for the identical question. A golden test would fail every single run, for no real reason, and everyone would learn to ignore it. You need a different kind of check: not "is this the exact expected text," but "does this output obey the rules a correct answer must obey."

How it works

   deterministic code               nondeterministic LLM output
   ------------------                ---------------------------
   same input                        same input
   -> ALWAYS the same output         -> a DIFFERENT output each time
   -> compare to one saved answer    -> check RULES the output must obey:

                                         - is it valid JSON, with the right fields?
                                         - is the length sensible?
                                         - does it avoid banned words?
                                         - does it stay on topic?

A real example you have seen

A customer-support chatbot rarely replies with the exact same sentence twice, even to the same question — and that is fine. What is not fine is a reply that is empty, that promises something the company cannot deliver, or that answers a completely different question. Those are the rules worth testing.

Remember this

  • When output is nondeterministic, test rules the output must obey, not one exact expected answer.
  • Run the check many times, with different random outcomes, because one lucky pass proves nothing.
  • Structural rules — valid JSON, the right fields, a sensible length — are the cheapest and most reliable checks to start with.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install pytest

No API key needed for this lesson. The examples below use a small stand-in for a real model call — it picks randomly from a few templates, the same way sampling at a temperature above zero does, but runs offline in milliseconds. Swap reply() and reply_json() for your real API call; the tests do not change shape when you do.

fake_llm.py
"""A stand-in for a real LLM call: randomised, like sampling at temperature > 0,
but running fully offline so the lesson is CPU-runnable.
"""
import json
import random

OPENERS = ["Sure!", "Absolutely.", "Here you go —", "Happy to help."]
CLOSERS = ["Let me know if that helps.", "Hope that's useful!", "Anything else?"]


def reply(question: str, seed: int | None = None) -> str:
    rng = random.Random(seed)
    opener = rng.choice(OPENERS)
    closer = rng.choice(CLOSERS)
    return f"{opener} The answer to '{question}' involves a few steps. {closer}"


def reply_json(question: str, seed: int | None = None) -> str:
    rng = random.Random(seed)
    payload = {
        "answer": f"a short answer about {question}",
        "confidence": round(rng.uniform(0.6, 0.99), 2),
        "sources": rng.randint(1, 4),
    }
    return json.dumps(payload)
python
from fake_llm import reply
for s in [1, 2, 3]:
    print(reply("how do I reset my password", seed=s))
Output
Absolutely. The answer to 'how do I reset my password' involves a few steps. Anything else?
Sure! The answer to 'how do I reset my password' involves a few steps. Let me know if that helps.
Absolutely. The answer to 'how do I reset my password' involves a few steps. Anything else?

Three different seeds, three different sentences. This is honest randomness, not a bug — a real model at a non-zero temperature behaves the same way, and no test should assume otherwise.

The tests

test_llm_output.py
import json

import pytest

from fake_llm import reply, reply_json

BANNED_WORDS = {"guarantee", "promise", "definitely"}


# --- Structural checks: shape and type, never exact text ---

def test_json_reply_has_every_required_field():
    data = json.loads(reply_json("refund policy", seed=1))
    assert set(data.keys()) == {"answer", "confidence", "sources"}


def test_confidence_is_always_a_valid_probability():
    for seed in range(20):
        data = json.loads(reply_json("refund policy", seed=seed))
        assert isinstance(data["confidence"], float)
        assert 0.0 <= data["confidence"] <= 1.0


def test_sources_is_always_a_positive_integer():
    for seed in range(20):
        data = json.loads(reply_json("refund policy", seed=seed))
        assert isinstance(data["sources"], int)
        assert data["sources"] >= 1


# --- Property checks: things true of every sample, not one fixed sample ---

def test_reply_is_never_empty_and_never_absurdly_long():
    for seed in range(20):
        text = reply("how do I reset my password", seed=seed)
        assert 10 <= len(text) <= 400


def test_reply_never_contains_a_banned_word():
    for seed in range(20):
        text = reply("how do I reset my password", seed=seed).lower()
        assert not any(word in text for word in BANNED_WORDS)


def test_reply_stays_on_topic():
    # a weak but cheap groundedness check: the question itself should be echoed
    for seed in range(20):
        text = reply("refund policy", seed=seed)
        assert "refund policy" in text
bash
pytest test_llm_output.py -q
Output
......                                                                   [100%]
6 passed in 0.07s

Line-by-line: why loop over 20 seeds

Every property test above loops over a range of seeds instead of checking one call. A single lucky sample can pass by accident. Looping over twenty different random outcomes checks the property holds across the distribution of things the model can say, not only the one thing it happened to say this time. This is the direct nondeterministic equivalent of the boundary-value testing from unit testing machine learning code — there you tested specific values on purpose; here you sample many values on purpose, because you cannot list them all by hand.

Now break something on purpose

Someone renames the confidence field to score while updating the prompt template — a real, common change when a prompt or schema gets "cleaned up":

python
# in fake_llm.py, reply_json becomes:
payload = {
    "answer": f"a short answer about {question}",
    "score": round(rng.uniform(0.6, 0.99), 2),   # renamed from "confidence"
    "sources": rng.randint(1, 4),
}
Output
FF....                                                                   [100%]
================================== FAILURES ===================================
__________________ test_json_reply_has_every_required_field ___________________

    def test_json_reply_has_every_required_field():
        data = json.loads(reply_json("refund policy", seed=1))
>       assert set(data.keys()) == {"answer", "confidence", "sources"}
E       AssertionError: assert {'answer', 'score', 'sources'} == {'answer', 'c...e', 'sources'}
E         
E         Extra items in the left set:
E         'score'
E         Extra items in the right set:
E         'confidence'
E         Use -v to get more diff

test_llm_output.py:14: AssertionError
________________ test_confidence_is_always_a_valid_probability ________________

    def test_confidence_is_always_a_valid_probability():
        for seed in range(20):
            data = json.loads(reply_json("refund policy", seed=seed))
>           assert isinstance(data["confidence"], float)
E           KeyError: 'confidence'

test_llm_output.py:20: KeyError
=========================== short test summary info ===========================
FAILED test_llm_output.py::test_json_reply_has_every_required_field - Asserti...
FAILED test_llm_output.py::test_confidence_is_always_a_valid_probability - Ke...
2 failed, 4 passed in 0.28s

Even though the content of every reply is different on every run, the structural tests catch the schema change instantly and exactly, in two precisely named failures. This is the real value of structural checks: they are strict about shape while staying completely indifferent to wording.

Common mistakes

Asserting on exact wording. assert text == "Sure! The answer..." will fail almost immediately, for no real reason, on the very next run. It is testing a property the system was never designed to have.

Testing with one seed and calling it done. One pass proves the property held once. It says nothing about the other 99% of what the model can produce. Loop across many samples, as above.

No temperature-zero option checked anywhere. Most real APIs let you request temperature=0, which makes output far more repeatable — not perfectly bit-identical across model versions, but close. Use it for tests that genuinely need repeatability, and reserve property-based tests for the parts that must handle real variation.

Skipping structural checks because "the model is usually good at JSON". Usually is not always. A schema check costs a few milliseconds and catches malformed output before it reaches a parser that was not written to handle it.

Try it yourself

Add a test that calls reply() a hundred times with different seeds and asserts that at least two distinct strings appear. This checks the opposite thing from every test above — not that variation stays within limits, but that variation genuinely exists, which would catch a stand-in (or a real API misconfigured to always return a cached answer) that stopped actually sampling.

What to learn next

Researcher — Mathematics and papers.

Why exact-match testing is the wrong tool here

Sampling from a language model draws a token sequence from a conditional distribution $p(y \mid x)$ over possible outputs $y$ given prompt $x$. For any non-degenerate distribution (temperature $> 0$, or top-$p$ / top-$k$ sampling with more than one candidate), $P(y = y_1) < 1$ for any specific $y_1$, so an exact-match test asserts a near-zero-probability event and will fail on almost every run by construction. The correct target is a property $\phi(y)$ that holds with probability approaching 1 over the support of $p(y \mid x)$ — schema validity, length bounds, content constraints — which is exactly what the structural and property tests above check.

Statistical testing of a stochastic system

Looping over $n$ seeds and asserting the property on every one of the $n$ samples is a weak form of hypothesis testing: it verifies $P(\phi(y)) \approx 1$ only to the resolution the sample size allows. If the true failure rate is $p$, the probability of seeing zero failures in $n$ independent samples is $(1-p)^n$. For $n = 20$ and a true failure rate of even 5%, there is a $(0.95)^{20} \approx 36\%$ chance the test suite reports all-green while the model fails on one call in twenty in production. Sizing $n$ against an acceptable failure rate is a direct application of the reasoning in statistical power and sample size; a CI budget of hundreds rather than tens of samples is common for anything shipping to real users.

Beyond structural checks: semantic and judged evaluation

Structural and rule-based checks (this lesson) verify form. They cannot verify that an answer is factually correct or genuinely relevant to the question — that requires either a reference answer compared by semantic similarity (embedding cosine similarity against a known-good response) or an LLM-as-judge, a second model prompted to score the first model's output against a rubric. Both introduce their own noise and cost, and belong at a later CI stage than the cheap structural checks in this lesson, gated in the way described in evaluation gates in CI.

Papers

  • Kwiatkowski et al. and the wider structured-generation literature — constrained decoding (grammar-restricted sampling) is the production-side complement to this lesson's test-side approach, guaranteeing schema validity at generation time rather than only checking it afterward.
  • Ribeiro et al., Beyond Accuracy: Behavioral Testing of NLP Models with CheckList, ACL 2020 — the property-style testing philosophy this lesson applies to a stochastic generator specifically.
  • Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, NeurIPS 2023 — the reliability and biases of using a model to grade another model's free-text output.

What to learn next