Latency, Load Testing and Capacity

Load testing a model endpoint

Load testing sends real, concurrent traffic at a service before real users do, so you find its breaking point on your own schedule instead of during a launch.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Load testing sends real, concurrent traffic at a service on purpose, so you find its limits before real users do.

The analogy you have already lived

You have tried squeezing one more person into a packed lift to see if the buzzer goes off. That is a deliberate test of a limit, done on your own terms, in a moment that does not matter. It beats finding out the hard way during an actual emergency.

Load testing is that same idea, aimed at a model server instead of a lift.

Why it exists

A server that answers one request nicely tells you almost nothing about what happens under real traffic. Ten people hitting it at once behaves differently from one person hitting it ten times in a row. Shared resources — a database connection, a worker pool, a single GPU — start to compete with each other.

Load testing means firing many requests at a service at the same time, on purpose. It measures how fast the service answers, how many requests fail, and at what point it starts to fall over.

Finding that point during a planned test, at 2pm on a Tuesday, is a very different experience from finding it during a product launch.

How it works

   load generator                    your server
   (many workers)
        |
        |--- request 1 --->
        |--- request 2 --->    all arriving
        |--- request 3 --->    close together
        |--- request 4 --->
        |
        <--- answers, one by one, as the server finishes each ---
        |
   measures: how many finished, how fast, how many failed

A real example you have seen

Before a big sale, e-commerce and ticket-booking sites run rehearsals with simulated traffic, precisely to find weak points before real shoppers arrive. A site that survives a sudden crowd without crashing was almost certainly load-tested first.

The honest part

A load test only tells you about the traffic pattern you actually sent. A steady stream of identical requests is not the same as a sudden spike. Neither is the same as real users doing a mix of different things. Test the shape of traffic you actually expect, not whatever is easiest to generate.

Remember this

  • Load testing means sending real, concurrent traffic on purpose, to find limits before users find them for you.
  • Shared resources — a database, a GPU, a worker pool — behave differently under real concurrency than under one request at a time.
  • A load test is only as useful as how closely it matches your real traffic pattern.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install fastapi "uvicorn[standard]" httpx

A tiny model server

Save this as server.py and run it with uvicorn server:app --host 127.0.0.1 --port 8000 in one terminal.

server.py
import random
import time
from fastapi import FastAPI

app = FastAPI()

@app.get("/predict")
def predict(x: float = 1.0):
    # Stand-in "model": a little real work plus jitter, so the timings
    # a client sees are genuine round trips, not invented numbers.
    time.sleep(max(0.001, random.gauss(0.008, 0.003)))
    return {"y": x * 2}

The load generator

Save this as load_test_client.py and run it in a second terminal, while the server from above is running.

load_test_client.py
import time
from concurrent.futures import ThreadPoolExecutor

import httpx

URL = "http://127.0.0.1:8000/predict"
N_REQUESTS = 300
CONCURRENCY = 20


def one_request(client):
    t0 = time.perf_counter()
    r = client.get(URL, params={"x": 3})
    return time.perf_counter() - t0, r.status_code


def percentile(values, p):
    s = sorted(values)
    return s[min(int(len(s) * p), len(s) - 1)]


if __name__ == "__main__":
    latencies, errors = [], 0
    start = time.perf_counter()
    with httpx.Client() as client, ThreadPoolExecutor(max_workers=CONCURRENCY) as pool:
        for elapsed, status in pool.map(lambda _: one_request(client), range(N_REQUESTS)):
            if status == 200:
                latencies.append(elapsed)
            else:
                errors += 1
    wall = time.perf_counter() - start

    print(f"{N_REQUESTS} requests, concurrency {CONCURRENCY}")
    print(f"wall time   : {wall:.2f}s")
    print(f"throughput  : {N_REQUESTS/wall:.1f} req/s")
    print(f"errors      : {errors}")
    print(f"p50 latency : {percentile(latencies, 0.50)*1000:6.1f} ms")
    print(f"p95 latency : {percentile(latencies, 0.95)*1000:6.1f} ms")
    print(f"p99 latency : {percentile(latencies, 0.99)*1000:6.1f} ms")
    print(f"worst       : {max(latencies)*1000:6.1f} ms")
Output
300 requests, concurrency 20
wall time   : 0.30s
throughput  : 995.4 req/s
errors      : 0
p50 latency :   12.4 ms
p95 latency :   18.7 ms
p99 latency :   20.2 ms
worst       :   30.8 ms

This is a real end-to-end measurement: a real HTTP server, a real network connection over localhost, and a real thread pool sending real concurrent requests. Rerunning it gives slightly different numbers — this is genuine timing over a real socket, and your machine's numbers will differ from these. The pattern — throughput near a thousand requests per second, p99 not far past p50 — is what matters here, not these exact figures.

Line-by-line walkthrough

ThreadPoolExecutor(max_workers=CONCURRENCY) is what makes this a load test rather than a sequential one. Twenty worker threads issue requests at once, because httpx releases the GIL while it waits on the network — see concurrency in a Python inference server for why that matters.

pool.map runs one_request across all 300 request indices, using the thread pool, and yields results back in order as they complete.

Common mistakes

Testing from the same machine the server runs on. The load generator competes with the server for CPU, and your numbers measure both at once, tangled together. Run them on separate machines when you can.

Only testing steady load. Real traffic spikes. Test a sudden burst — all requests fired at once, instead of spread evenly — since it stresses the queue and worker pool differently.

Ignoring errors in the summary. A test that reports latency for successful requests only, while quietly dropping the failed ones, can make an overloaded server look healthy. Always report the error count next to the latency.

Not warming the server up first. The first few requests to a freshly started server are often slower — connection pools filling, caches cold. Send a short warm-up burst before the real measurement starts.

Try it yourself

Change CONCURRENCY from 20 to 100 and rerun. Watch p99 and worst — they should grow much faster than p50 does, since the server's real bottleneck starts to show only under heavier concurrency.

What to learn next

Researcher — Mathematics and papers.

Open-loop versus closed-loop generation

The load generator above is closed-loop: each worker only issues its next request after its previous one completes, so total offered load is bounded by CONCURRENCY regardless of how slow the server gets. An open-loop generator instead issues requests on a fixed schedule, independent of whether earlier requests have returned — a far closer model of real, independent users, and the only design that avoids coordinated omission, covered next.

What a load test needs to vary deliberately

  • Concurrency — how many requests are in flight at once.
  • Arrival pattern — steady rate, Poisson arrivals, or a deliberate burst.
  • Payload shape — realistic input sizes; a model server's cost is rarely flat across input sizes.
  • Duration — long enough to expose problems that only appear after warm-up, cache fill, or memory growth (see soak testing).

Tooling beyond a hand-rolled script

Purpose-built load testers solve problems a script like the one above does not: k6 and Gatling generate load from a compiled, low-overhead engine so the tool itself is not the bottleneck; Locust distributes load generation across multiple machines; wrk2 and tools built on HdrHistogram specifically correct for coordinated omission by scheduling requests against wall-clock time rather than against the previous response. For anything beyond a quick local check, reaching for one of these is usually better than extending a hand-rolled script indefinitely.

Statistical validity of a single run

A single load test run is one sample from a noisy process — background load on shared infrastructure, JIT or cache warm-up effects, and network jitter all move the numbers. Reporting a load test's result as a single number without a confidence interval or multiple repeated runs is a common and avoidable source of false conclusions, especially when comparing two versions of a service against each other.

References

  • Tene, G., How NOT to Measure Latency, Strange Loop 2015 — the canonical talk on closed-loop measurement pitfalls in load testing, expanded on in the next lesson.
  • k6 documentation, Test types — a practical taxonomy of load, stress, soak and spike testing.

What to learn next