Testing ML Code and CI

Integration testing an inference server

An integration test starts the real server as its own process and talks to it over real HTTP, the way a health inspector checks a restaurant by walking in the front door instead of reading the kitchen's recipe cards.

On this page 7
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

An integration test starts your real server as its own process and talks to it over real HTTP, exactly like a real caller would.

The analogy you have already lived

A restaurant health inspector does not read the kitchen's recipe cards and call it a check. They walk in the front door, order food like a customer, and see what actually arrives at the table.

Unit tests are the recipe cards — they check one function at a time, on its own. An integration test is the inspector walking in. It starts the whole server, sends it a real request the way a real caller would, and checks what comes back.

Why it exists

A service can pass every unit test and still fail the moment it actually starts. The startup code might have a typo in a file path. Two parts that work fine alone might disagree about the shape of data passed between them. A setting that looks fine in a config file might be wrong for the actual machine it runs on.

None of that shows up in a unit test, because a unit test never starts the real process. An integration test does exactly that. Only that class of test can catch a service broken specifically at the seams between its parts. The same is true for a service broken only when it truly starts up for real.

How it works

   test code
       |
       | starts the REAL server as its own process
       v
   [ your server, running for real, on a real port ]
       ^
       |  real HTTP requests, over the network stack, like any caller
       |
   test code checks:
       - does /healthz say "ready" before we send real traffic?
       - does a normal request get a normal answer?
       - does a bad request get a clean error, not a crash?
       - is the server still alive after a burst of requests?

A real example you have seen

A UPI app can work perfectly in a demo. Then it fails the first time it is opened on a real phone, on a real network, with a real bank's server on the other end. That is exactly the gap integration testing is built to close: everything worked in isolation, and something broke where the real pieces met.

Remember this

  • An integration test runs the real, whole service, not one function in isolation.
  • It catches problems at the seams — startup, networking, timing — that unit tests structurally cannot see.
  • Checking the server is still alive after the test is as important as checking the answer was correct.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install fastapi "uvicorn[standard]" scikit-learn pandas joblib requests pytest

This continues from model serving — train.py and serve.py are the same loan-scorer service built there. Run train.py once before the tests.

train.py
import joblib
import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

rng = np.random.RandomState(0)
n = 400
X = pd.DataFrame({
    "income": rng.uniform(5, 80, n).round(1),
    "years": rng.uniform(0, 10, n).round(1),
    "age": rng.randint(21, 65, n).astype(float),
})
score = 0.05 * X["income"] + 0.35 * X["years"] + 0.01 * X["age"] - 3.5
y = (score + rng.normal(0, 0.8, n) > 0).astype(int)

model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000)).fit(X, y)
joblib.dump(model, "model.joblib")
serve.py
from contextlib import asynccontextmanager

import joblib
import pandas as pd
from fastapi import FastAPI
from pydantic import BaseModel, Field

STATE = {}


@asynccontextmanager
async def lifespan(app: FastAPI):
    STATE["model"] = joblib.load("model.joblib")
    STATE["columns"] = list(STATE["model"].feature_names_in_)
    yield
    STATE.clear()


app = FastAPI(title="loan-scorer", lifespan=lifespan)


class Applicant(BaseModel):
    income: float = Field(ge=0, le=1000)
    years: float = Field(ge=0, le=60)
    age: float = Field(ge=18, le=100)


@app.get("/healthz")
def healthz():
    return {"status": "ok", "columns": STATE.get("columns")}


@app.post("/predict")
def predict(applicant: Applicant):
    row = pd.DataFrame([applicant.model_dump()])[STATE["columns"]]
    probability = float(STATE["model"].predict_proba(row)[0, 1])
    return {"repaid": int(probability >= 0.5), "probability": round(probability, 4)}

The integration test

This starts the server as a genuinely separate process and talks to it over a real socket — the thing model serving's in-process TestClient deliberately avoids doing, because that keeps unit-style tests fast. This lesson trades that speed for a stronger guarantee: the exact same startup path a real deployment uses.

test_integration.py
import subprocess
import sys
import time

import pytest
import requests

BASE_URL = "http://127.0.0.1:8931"


@pytest.fixture(scope="module")
def running_server():
    """Starts the real server as a separate process -- a genuine black-box test."""
    proc = subprocess.Popen(
        [sys.executable, "-m", "uvicorn", "serve:app", "--port", "8931"],
        stdout=subprocess.PIPE, stderr=subprocess.STDOUT, text=True,
    )
    deadline = time.time() + 10
    while time.time() < deadline:
        try:
            r = requests.get(f"{BASE_URL}/healthz", timeout=0.5)
            if r.status_code == 200:
                break
        except requests.ConnectionError:
            time.sleep(0.2)
    else:
        proc.terminate()
        raise RuntimeError("server did not become healthy in time:\n" + proc.stdout.read())

    yield proc

    proc.terminate()
    proc.wait(timeout=5)


def test_health_check_reports_the_model_is_loaded(running_server):
    r = requests.get(f"{BASE_URL}/healthz")
    assert r.status_code == 200
    assert r.json()["columns"] == ["income", "years", "age"]


def test_a_real_http_prediction_round_trip(running_server):
    r = requests.post(f"{BASE_URL}/predict", json={"income": 60.0, "years": 7.0, "age": 34})
    assert r.status_code == 200
    body = r.json()
    assert body["repaid"] in (0, 1)
    assert 0.0 <= body["probability"] <= 1.0


def test_a_malformed_request_gets_a_422_not_a_crash(running_server):
    r = requests.post(f"{BASE_URL}/predict", json={"income": 60.0})
    assert r.status_code == 422
    assert running_server.poll() is None  # the server process is still alive


def test_the_server_survives_a_burst_of_requests(running_server):
    results = [
        requests.post(f"{BASE_URL}/predict", json={"income": i, "years": 3.0, "age": 30})
        for i in range(20)
    ]
    assert all(r.status_code == 200 for r in results)
bash
pytest test_integration.py -q
Output
....                                                                     [100%]
4 passed in 3.08s

Line-by-line: the polling loop

The running_server fixture does not assume the server is instantly ready. It polls /healthz in a loop, up to a ten-second deadline, catching the connection error that happens in the window between "process started" and "port is actually accepting connections". This detail is easy to skip and is exactly why "flaky" integration tests exist — a fixed time.sleep(1) instead of a real poll works on a fast machine and fails at random on a loaded one.

test_a_malformed_request_gets_a_422_not_a_crash checks two separate things: the right status code, and that running_server.poll() still returns None — meaning the process has not exited. A server that crashes on bad input and returns a correct-looking error to that one request, while dying for every request after it, is a real and dangerous failure mode this second assertion is built to catch.

Now break something on purpose

Delete model.joblib before starting the server, simulating a real deployment mistake — the model artifact was never copied into the image:

bash
mv model.joblib model.joblib.bak
pytest test_integration.py::test_health_check_reports_the_model_is_loaded -q
Output
RuntimeError: server did not become healthy in time:
INFO:     Started server process [67240]
INFO:     Waiting for application startup.
ERROR:    Traceback (most recent call last):
  File ".../starlette/routing.py", line 648, in lifespan
    async with self.lifespan_context(app) as maybe_state:
  File ".../serve.py", line 13, in lifespan
    STATE["model"] = joblib.load("model.joblib")
FileNotFoundError: [Errno 2] No such file or directory: 'model.joblib'
ERROR:    Application startup failed. Exiting.

The whole real startup log is captured and shown to you, because stderr=subprocess.STDOUT merged it into the same stream the fixture reads on timeout. This is a genuine ten-second wait — the fixture polls for the full deadline before giving up, which is the honest cost of testing a real startup sequence rather than a mocked one. It is also exactly the failure a bad deploy produces in production, caught here before it ever reaches one.

Common mistakes

Hardcoding a fixed sleep instead of polling. time.sleep(2) works until the CI machine is slower than your laptop one day, and then it does not. Poll a real endpoint until it answers.

Never checking the process is still alive. A response with the right status code from a server that is about to crash is still a passing assertion, and a false sense of safety.

Running integration tests as often as unit tests. They are slower by nature — starting a real process, waiting for a real port. Run the fast unit and behavioural suites on every save; run this suite in CI on every push, not on every keystroke.

Forgetting to terminate the server. A test suite that starts subprocesses and never cleans them up leaves orphaned servers holding a port open, which then makes the next test run fail with "address already in use" for a completely unrelated reason.

Testing only the happy path. The malformed-request and burst-of-requests tests above exist because a server that only works when everything goes right is not tested — it is demonstrated.

Try it yourself

Add a test that sends ten requests concurrently, using concurrent.futures.ThreadPoolExecutor, instead of one after another. This checks something the sequential burst test above cannot: whether the server handles requests arriving at the same moment, which is closer to real traffic than a request queue you control yourself.

What to learn next

Researcher — Mathematics and papers.

Where this sits in the test pyramid

The conventional test pyramid places unit tests at the base (many, fast, isolated), integration tests in the middle (fewer, slower, real collaborators), and end-to-end tests at the top (fewest, slowest, full user-facing flow). This lesson's tests sit at the integration layer specifically: real process, real network stack, but a single service — no upstream dependencies, no browser, no real database. Cohn's original formulation (2009) argues for a roughly 70/20/10 split by count across the three layers; ML services in practice often skew integration-heavy relative to that guideline, because so many real failures (as demonstrated above) occur specifically at startup and at the network boundary rather than inside pure functions.

What subprocess isolation buys you over in-process testing

FastAPI's TestClient (used in model serving) runs the ASGI application directly in the test process, using the ASGI protocol rather than real sockets. This is faster and sufficient for testing request handling logic, but it shares the test process's Python interpreter, environment variables, and import state with the application under test — meaning a bug that only manifests under a genuinely separate process (an environment variable the test happened to set, a module-level side effect from another test's import) is invisible to it. Subprocess-based integration testing removes that shared state entirely, at the cost of the process-startup latency measured above.

Readiness versus liveness

The polling loop against /healthz implements a readiness check: is the service ready to receive traffic? This is a distinct concept from a liveness check: is the process still running at all, used by an orchestrator to decide whether to restart a container. Kubernetes formalises this split explicitly (readinessProbe vs livenessProbe); conflating them is a common production bug, where a service that is alive but still loading a large model gets traffic routed to it anyway and returns errors until it finishes loading — the exact race this lesson's polling loop is designed to make visible in a test rather than in production.

Papers and prior art

  • Cohn, Succeeding with Agile, 2009 — the test pyramid this lesson's place in the suite is drawn from.
  • Fowler, TestPyramid, martinfowler.com, 2012 — martinfowler.com/bliki/TestPyramid.html
  • Breck et al., The ML Test Score, IEEE Big Data 2017 — Infra Test 5 ("the model is tested via its serving API") is precisely the property this lesson's suite establishes.

What to learn next