Testing output you cannot predict
When a model's exact wording changes every time, you test the shape and the rules the output must always obey — the way you judge two cups of chai by heat and sweetness, not by counting identical tea leaves.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
When exact output cannot be predicted, you test the rules it must always obey, instead of one exact expected answer.
The analogy you have already lived
You have ordered chai from the same stall twice. The two cups were not identical — a little more or less sugar, poured slightly differently. You did not complain about that. You would complain if one arrived cold, or without any tea in it at all.
You were testing properties — hot, sweet, actually tea — not testing for one exact liquid. Large language models are the same. Ask the same question twice and the wording often differs, because these models pick from many likely next words rather than one fixed answer.
Why it exists
The golden output test from the previous lesson checks that an output matches a saved one, exactly. That works for a deterministic function — same input, same output, always.
A model that samples its next words does not give the same output twice, even for the identical question. A golden test would fail every single run, for no real reason, and everyone would learn to ignore it. You need a different kind of check: not "is this the exact expected text," but "does this output obey the rules a correct answer must obey."
How it works
deterministic code nondeterministic LLM output
------------------ ---------------------------
same input same input
-> ALWAYS the same output -> a DIFFERENT output each time
-> compare to one saved answer -> check RULES the output must obey:
- is it valid JSON, with the right fields?
- is the length sensible?
- does it avoid banned words?
- does it stay on topic?A real example you have seen
A customer-support chatbot rarely replies with the exact same sentence twice, even to the same question — and that is fine. What is not fine is a reply that is empty, that promises something the company cannot deliver, or that answers a completely different question. Those are the rules worth testing.
Remember this
- When output is nondeterministic, test rules the output must obey, not one exact expected answer.
- Run the check many times, with different random outcomes, because one lucky pass proves nothing.
- Structural rules — valid JSON, the right fields, a sensible length — are the cheapest and most reliable checks to start with.
What to learn next
- Recording and replaying model API calls — testing code that calls a model without paying for or depending on the real API every run.
- Evaluation gates in CI — turning checks like these into a pass/fail line that blocks a bad release.
- LLM-as-judge in production — using a second model to grade answers a rule cannot check.
Developer — Code and libraries.
Setup
pip install pytestNo API key needed for this lesson. The examples below use a small stand-in for a real model call — it picks randomly from a few templates, the same way sampling at a temperature above zero does, but runs offline in milliseconds. Swap reply() and reply_json() for your real API call; the tests do not change shape when you do.
"""A stand-in for a real LLM call: randomised, like sampling at temperature > 0,
but running fully offline so the lesson is CPU-runnable.
"""
import json
import random
OPENERS = ["Sure!", "Absolutely.", "Here you go —", "Happy to help."]
CLOSERS = ["Let me know if that helps.", "Hope that's useful!", "Anything else?"]
def reply(question: str, seed: int | None = None) -> str:
rng = random.Random(seed)
opener = rng.choice(OPENERS)
closer = rng.choice(CLOSERS)
return f"{opener} The answer to '{question}' involves a few steps. {closer}"
def reply_json(question: str, seed: int | None = None) -> str:
rng = random.Random(seed)
payload = {
"answer": f"a short answer about {question}",
"confidence": round(rng.uniform(0.6, 0.99), 2),
"sources": rng.randint(1, 4),
}
return json.dumps(payload)from fake_llm import reply
for s in [1, 2, 3]:
print(reply("how do I reset my password", seed=s))Absolutely. The answer to 'how do I reset my password' involves a few steps. Anything else? Sure! The answer to 'how do I reset my password' involves a few steps. Let me know if that helps. Absolutely. The answer to 'how do I reset my password' involves a few steps. Anything else?
Three different seeds, three different sentences. This is honest randomness, not a bug — a real model at a non-zero temperature behaves the same way, and no test should assume otherwise.
The tests
import json
import pytest
from fake_llm import reply, reply_json
BANNED_WORDS = {"guarantee", "promise", "definitely"}
# --- Structural checks: shape and type, never exact text ---
def test_json_reply_has_every_required_field():
data = json.loads(reply_json("refund policy", seed=1))
assert set(data.keys()) == {"answer", "confidence", "sources"}
def test_confidence_is_always_a_valid_probability():
for seed in range(20):
data = json.loads(reply_json("refund policy", seed=seed))
assert isinstance(data["confidence"], float)
assert 0.0 <= data["confidence"] <= 1.0
def test_sources_is_always_a_positive_integer():
for seed in range(20):
data = json.loads(reply_json("refund policy", seed=seed))
assert isinstance(data["sources"], int)
assert data["sources"] >= 1
# --- Property checks: things true of every sample, not one fixed sample ---
def test_reply_is_never_empty_and_never_absurdly_long():
for seed in range(20):
text = reply("how do I reset my password", seed=seed)
assert 10 <= len(text) <= 400
def test_reply_never_contains_a_banned_word():
for seed in range(20):
text = reply("how do I reset my password", seed=seed).lower()
assert not any(word in text for word in BANNED_WORDS)
def test_reply_stays_on_topic():
# a weak but cheap groundedness check: the question itself should be echoed
for seed in range(20):
text = reply("refund policy", seed=seed)
assert "refund policy" in textpytest test_llm_output.py -q...... [100%] 6 passed in 0.07s
Line-by-line: why loop over 20 seeds
Every property test above loops over a range of seeds instead of checking one call. A single lucky sample can pass by accident. Looping over twenty different random outcomes checks the property holds across the distribution of things the model can say, not only the one thing it happened to say this time. This is the direct nondeterministic equivalent of the boundary-value testing from unit testing machine learning code — there you tested specific values on purpose; here you sample many values on purpose, because you cannot list them all by hand.
Now break something on purpose
Someone renames the confidence field to score while updating the prompt template — a real, common change when a prompt or schema gets "cleaned up":
# in fake_llm.py, reply_json becomes:
payload = {
"answer": f"a short answer about {question}",
"score": round(rng.uniform(0.6, 0.99), 2), # renamed from "confidence"
"sources": rng.randint(1, 4),
}FF.... [100%]
================================== FAILURES ===================================
__________________ test_json_reply_has_every_required_field ___________________
def test_json_reply_has_every_required_field():
data = json.loads(reply_json("refund policy", seed=1))
> assert set(data.keys()) == {"answer", "confidence", "sources"}
E AssertionError: assert {'answer', 'score', 'sources'} == {'answer', 'c...e', 'sources'}
E
E Extra items in the left set:
E 'score'
E Extra items in the right set:
E 'confidence'
E Use -v to get more diff
test_llm_output.py:14: AssertionError
________________ test_confidence_is_always_a_valid_probability ________________
def test_confidence_is_always_a_valid_probability():
for seed in range(20):
data = json.loads(reply_json("refund policy", seed=seed))
> assert isinstance(data["confidence"], float)
E KeyError: 'confidence'
test_llm_output.py:20: KeyError
=========================== short test summary info ===========================
FAILED test_llm_output.py::test_json_reply_has_every_required_field - Asserti...
FAILED test_llm_output.py::test_confidence_is_always_a_valid_probability - Ke...
2 failed, 4 passed in 0.28sEven though the content of every reply is different on every run, the structural tests catch the schema change instantly and exactly, in two precisely named failures. This is the real value of structural checks: they are strict about shape while staying completely indifferent to wording.
Common mistakes
Asserting on exact wording. assert text == "Sure! The answer..." will fail almost immediately, for no real reason, on the very next run. It is testing a property the system was never designed to have.
Testing with one seed and calling it done. One pass proves the property held once. It says nothing about the other 99% of what the model can produce. Loop across many samples, as above.
No temperature-zero option checked anywhere. Most real APIs let you request temperature=0, which makes output far more repeatable — not perfectly bit-identical across model versions, but close. Use it for tests that genuinely need repeatability, and reserve property-based tests for the parts that must handle real variation.
Skipping structural checks because "the model is usually good at JSON". Usually is not always. A schema check costs a few milliseconds and catches malformed output before it reaches a parser that was not written to handle it.
Try it yourself
Add a test that calls reply() a hundred times with different seeds and asserts that at least two distinct strings appear. This checks the opposite thing from every test above — not that variation stays within limits, but that variation genuinely exists, which would catch a stand-in (or a real API misconfigured to always return a cached answer) that stopped actually sampling.
What to learn next
- Recording and replaying model API calls — testing code that calls a model without paying for or depending on the real API every run.
- Evaluation gates in CI — turning checks like these into a pass/fail line that blocks a bad release.
- LLM-as-judge in production — using a second model to grade answers a rule cannot check.
Researcher — Mathematics and papers.
Why exact-match testing is the wrong tool here
Sampling from a language model draws a token sequence from a conditional distribution $p(y \mid x)$ over possible outputs $y$ given prompt $x$. For any non-degenerate distribution (temperature $> 0$, or top-$p$ / top-$k$ sampling with more than one candidate), $P(y = y_1) < 1$ for any specific $y_1$, so an exact-match test asserts a near-zero-probability event and will fail on almost every run by construction. The correct target is a property $\phi(y)$ that holds with probability approaching 1 over the support of $p(y \mid x)$ — schema validity, length bounds, content constraints — which is exactly what the structural and property tests above check.
Statistical testing of a stochastic system
Looping over $n$ seeds and asserting the property on every one of the $n$ samples is a weak form of hypothesis testing: it verifies $P(\phi(y)) \approx 1$ only to the resolution the sample size allows. If the true failure rate is $p$, the probability of seeing zero failures in $n$ independent samples is $(1-p)^n$. For $n = 20$ and a true failure rate of even 5%, there is a $(0.95)^{20} \approx 36\%$ chance the test suite reports all-green while the model fails on one call in twenty in production. Sizing $n$ against an acceptable failure rate is a direct application of the reasoning in statistical power and sample size; a CI budget of hundreds rather than tens of samples is common for anything shipping to real users.
Beyond structural checks: semantic and judged evaluation
Structural and rule-based checks (this lesson) verify form. They cannot verify that an answer is factually correct or genuinely relevant to the question — that requires either a reference answer compared by semantic similarity (embedding cosine similarity against a known-good response) or an LLM-as-judge, a second model prompted to score the first model's output against a rubric. Both introduce their own noise and cost, and belong at a later CI stage than the cheap structural checks in this lesson, gated in the way described in evaluation gates in CI.
Papers
- Kwiatkowski et al. and the wider structured-generation literature — constrained decoding (grammar-restricted sampling) is the production-side complement to this lesson's test-side approach, guaranteeing schema validity at generation time rather than only checking it afterward.
- Ribeiro et al., Beyond Accuracy: Behavioral Testing of NLP Models with CheckList, ACL 2020 — the property-style testing philosophy this lesson applies to a stochastic generator specifically.
- Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, NeurIPS 2023 — the reliability and biases of using a model to grade another model's free-text output.
What to learn next
- Recording and replaying model API calls — testing code that calls a model without paying for or depending on the real API every run.
- Evaluation gates in CI — turning checks like these into a pass/fail line that blocks a bad release.
- LLM-as-judge in production — using a second model to grade answers a rule cannot check.