Recording and replaying model API calls
A recorded cassette saves a real model call once and replays it in tests forever after, the way keeping a photocopy of a form saves you from standing in the same queue twice.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Recording and replaying means saving a real API response once, then reusing it in every test afterward instead of calling the real thing again.
The analogy you have already lived
You have photocopied a filled-in form before submitting it. Next time someone asks for the same information, you hand over the photocopy instead of standing in the same government-office queue again. The queue was real and slow the first time. The photocopy is instant, every time after.
A cassette — the common name for this pattern — is that photocopy for an API call. Call the real model once, save exactly what it said, and reuse the saved copy in every test run from then on.
Why it exists
Code that calls a large language model usually sits behind a paid API, reachable only over the network. Testing that code by calling the real API every time is slow, and costs money on every test run. It can also fail for reasons that have nothing to do with your code. The network is down, the API is rate-limiting you, a token expired.
Worse, the previous lesson showed that model output is not the same twice. A test that calls the real API and checks for exact wording will fail at random, even when your code is completely correct.
Recording the response once and replaying it removes all three problems. Your tests run offline, for free, in milliseconds, and check the same fixed response every time.
How it works
FIRST TIME (recording — a human runs this on purpose, once)
your code -> real API call -> real response -> SAVE to a cassette file
EVERY TEST RUN AFTER THAT (replay — this is what CI runs)
your code -> "was this exact request recorded?"
|
yes -> return the saved response, no network involved
|
no -> fail with a clear message: "record a cassette for this first"The cassette file is checked into version control, exactly like the golden files from the previous lesson. It is data your tests depend on, and it should be reviewed like any other change.
A real example you have seen
A recorded phone call played back at a call centre for training purposes is the same idea. The real conversation happened once, and everyone after that learns from the recording, without needing to place the real call again.
Remember this
- A cassette saves a real response once and replays it in every later test — no network, no cost, no randomness.
- Recording is a deliberate, human-run action, never something a test does automatically for itself.
- A request that was never recorded should fail with a clear message, telling you to record it — not silently call the real API.
What to learn next
- Integration testing an inference server — testing your own service end to end, the layer this lesson's mocking makes possible without a live upstream model.
- Golden outputs and regression tests — the same "save it once, compare forever" idea applied to deterministic code.
- Evaluation gates in CI — where recorded and live checks both feed into a single pass/fail decision.
Developer — Code and libraries.
Setup
pip install pytestReal projects usually reach for a library like vcrpy or respx to do this at the HTTP level. Building a small version yourself, with nothing but the standard library, makes the mechanism completely visible — and the idea transfers directly to whichever library you use later.
The real client
"""The real, network-calling side of the client. Never imported directly by tests."""
import hashlib
import random
import time
CALL_COUNT = 0 # only real calls increment this -- tests check it stays at 0
def call_llm_api(prompt: str) -> str:
"""Stands in for a real network call: costs time, costs money, and is not
guaranteed to return the same thing twice. Tests must never trigger this.
"""
global CALL_COUNT
CALL_COUNT += 1
time.sleep(0.01) # standing in for real network latency
# hashlib, not Python's built-in hash() -- hash() is randomised per
# process by design, which would make this "fake API" answer differently
# every time you ran it, defeating the entire point of a cassette.
seed = int(hashlib.md5(prompt.encode()).hexdigest()[:8], 16)
rng = random.Random(seed)
return f"Summary: {prompt[:20]}... ({rng.randint(1, 5)} key points)"The cassette layer
"""A minimal record/replay layer for LLM calls, using only the standard library."""
import json
import os
from functools import wraps
MODE = os.environ.get("CASSETTE_MODE", "replay") # "record" or "replay"
def cassette(path: str):
def decorator(real_fn):
@wraps(real_fn)
def wrapper(prompt: str) -> str:
tape = json.load(open(path)) if os.path.exists(path) else {}
if MODE == "record":
response = real_fn(prompt)
tape[prompt] = response
json.dump(tape, open(path, "w"), indent=2, sort_keys=True)
return response
if prompt not in tape:
raise KeyError(
f"No recorded response for prompt {prompt!r}. "
f"Re-record with CASSETTE_MODE=record."
)
return tape[prompt]
return wrapper
return decoratorThe code being tested
from cassette import cassette
from llm_client import call_llm_api
get_summary = cassette("cassette_summaries.json")(call_llm_api)
def summarize_ticket(ticket_text: str) -> str:
prompt = f"Summarize this support ticket: {ticket_text}"
return get_summary(prompt)Recording the cassette once
CASSETTE_MODE=record python3 -c "
from summarizer import summarize_ticket
print(summarize_ticket('My payment failed twice this week and I was charged both times.'))
print(summarize_ticket('How do I change my delivery address after ordering?'))
"Summary: Summarize this suppo... (5 key points) Summary: Summarize this suppo... (3 key points)
This writes cassette_summaries.json, mapping each exact prompt to the response that came back. That file gets committed to version control, next to the code.
The tests — run in replay mode, the CI default
import time
import pytest
import llm_client
from summarizer import summarize_ticket
def test_replay_never_calls_the_real_api():
before = llm_client.CALL_COUNT
result = summarize_ticket("My payment failed twice this week and I was charged both times.")
assert result.startswith("Summary:")
assert llm_client.CALL_COUNT == before # unchanged: the real function never ran
def test_replay_is_fast():
start = time.perf_counter()
for _ in range(50):
summarize_ticket("How do I change my delivery address after ordering?")
elapsed = time.perf_counter() - start
assert elapsed < 0.5 # 50 replayed calls should be nowhere near 50 real ones (0.01s each)
def test_an_unrecorded_prompt_fails_with_a_clear_error():
with pytest.raises(KeyError, match="No recorded response"):
summarize_ticket("This exact ticket was never recorded before.")pytest test_summarizer.py -q... [100%] 3 passed in 0.03s
Line-by-line: proving the real API never ran
test_replay_never_calls_the_real_api checks llm_client.CALL_COUNT, a counter that only increases inside the real call_llm_api function itself. If replay mode ever accidentally fell through to a real call, this counter would move, and the test would catch it precisely — not by guessing from timing, but by checking a direct side effect of the real function actually running.
test_replay_is_fast gives the same fact a second, independent kind of evidence: a real measurement, one run, on one machine.
one real call: 10.32 ms 50 replayed calls: 7.88 ms total, 0.1576 ms each
A single real call (with its simulated network delay) took about 10 ms. Fifty replayed calls together took about 8 ms — each one roughly sixty times faster than one real call. Your exact numbers will differ; the direction and the rough scale will not, because reading a small JSON file is inherently far faster than a network round trip.
Now break something on purpose
Code calling summarize_ticket with a prompt that was never recorded:
from summarizer import summarize_ticket
summarize_ticket("This exact ticket was never recorded before.")KeyError: "No recorded response for prompt 'Summarize this support ticket: This exact ticket was never recorded before.'. Re-record with CASSETTE_MODE=record."
This is the correct failure, not a bug. It means new test code needs a fresh recording — a clear, actionable message telling you exactly what to do, instead of a confusing network timeout or a silently wrong answer.
Common mistakes
Matching the cassette on more than the exact prompt. Real HTTP-level tools like vcrpy let you match loosely (ignore headers, ignore an API key), which is usually what you want — matching too strictly means a harmless change (like a timestamp in a header) breaks every recorded test.
Committing a cassette that contains a real secret. A recorded response can leak an API key if it was embedded in the request URL or headers rather than the body. Scrub cassettes before committing them, and check the specific library's redaction settings.
Never re-recording. A cassette recorded a year ago tests against a model version that may no longer exist. Re-record on a schedule, or whenever the underlying prompt or model changes on purpose — treat a stale cassette as a stale golden file.
Recording in CI by accident. If CASSETTE_MODE ever defaults to "record" in the CI environment, every test run silently calls the real API and pays for it. Keep "replay" as the hard default, as this lesson's code does, and require recording to be requested explicitly.
Try it yourself
Add a second cassette file for a different feature, and write a test that asserts a missing cassette file (not only a missing prompt) fails with a clear message rather than an unhandled FileNotFoundError. Real teams hit this exact failure the first time a new contributor checks out the repository without the cassette files.
What to learn next
- Integration testing an inference server — testing your own service end to end, the layer this lesson's mocking makes possible without a live upstream model.
- Golden outputs and regression tests — the same "save it once, compare forever" idea applied to deterministic code.
- Evaluation gates in CI — where recorded and live checks both feed into a single pass/fail decision.
Researcher — Mathematics and papers.
This pattern's place among test doubles
Fowler's (2007) taxonomy of test doubles — dummy, stub, spy, mock, fake — places a cassette closest to a stub with recorded state: it returns canned answers to specific calls, without asserting anything about how it was called (that would make it a mock in the strict sense) and without reimplementing real behaviour (that would make it a fake). The distinction matters because a stub-style cassette cannot verify how many times or in what order a dependency was called — only what it returned. Testing call order or count requires a genuine mock, layered on top if needed.
Where this pattern is standard practice
HTTP-level record/replay is well established outside ML: Ruby's VCR library (2008) coined the "cassette" name; Python's vcrpy and the newer respx (built for httpx) are direct descendants. LLM-specific tooling (LangSmith, Helicone, and similar observability platforms) increasingly builds cassette-equivalent replay directly into their tracing layer, blurring the line between a test fixture and a production observability trace — the same recorded request/response pair often serves both purposes.
The staleness problem, formally
A cassette encodes an implicit assumption: that $p_{\text{model}}(\cdot \mid \text{prompt})$ at recording time is a valid stand-in for $p_{\text{model}}(\cdot \mid \text{prompt})$ at test time. This assumption silently breaks whenever the provider updates the underlying model behind a stable model name — a documented, recurring failure mode for hosted APIs, distinct from the prompt-schema staleness a golden test catches. No purely offline test can detect this; it requires a periodic contract test that does call the real API on a schedule (nightly, not per-commit) and checks the response still matches the cassette's shape, if not its exact content.
Papers and prior art
- Fowler, Mocks Aren't Stubs, 2007 — martinfowler.com/articles/mocksArentStubs.html
- Freeman and Pryce, Growing Object-Oriented Software, Guided by Tests, 2009 — the broader case for isolating a system under test from slow, non-deterministic collaborators.
What to learn next
- Integration testing an inference server — testing your own service end to end, the layer this lesson's mocking makes possible without a live upstream model.
- Golden outputs and regression tests — the same "save it once, compare forever" idea applied to deterministic code.
- Evaluation gates in CI — where recorded and live checks both feed into a single pass/fail decision.