Latency, Load Testing and Capacity

Latency budgets

A latency budget splits your total allowed response time across every step of a request, so you know exactly which step is eating the most time before it becomes a problem.

Read these first

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A latency budget splits your total allowed response time across every step of a request.

The analogy you have already lived

You have planned a trip to catch a train. You give yourself thirty minutes total. Five to walk to the auto stand, fifteen for the ride, five to walk in, and five spare in case something goes wrong. Each leg gets its own slice of the thirty minutes.

If the auto ride runs long, you know exactly which leg to worry about, and how much of your spare time it ate.

A model server's response time works the same way, split across its own steps instead of legs of a journey.

Why it exists

A request rarely does only one thing. It checks who is asking, fetches some data, runs the model, and formats a reply. Each of those steps takes some time.

Add up every step's time, and you get the total. Suppose nobody ever wrote down how much time each step is allowed to take. Then nobody notices when one step quietly starts taking three times longer than it used to. Users start complaining before anyone else does.

A latency budget writes that allowance down in advance. Auth gets 3ms. Fetching data gets 8ms. The model gets 12ms. Formatting gets 1ms. Add a little slack, and you have a target for the whole request — and, more importantly, a target for each piece of it.

How it works

total budget: 30ms

  auth check        [===]                    3ms
  fetch features     [========]               8ms
  model call          [============]         12ms
  format response      [=]                    1ms
                                       -----------
                                       total: 24ms
                                       spare:  6ms

If fetch features starts taking 15ms instead of 8ms, the budget shows it immediately. The total no longer fits, and you know exactly where the extra time went.

A real example you have seen

A UPI payment confirming in under a second involves several checks happening one after another — your bank, the recipient's bank, a fraud check. Each has to fit inside a slice of that one second, or the whole payment feels slow, even if only one step was the actual problem.

The honest part

A budget that adds up on paper does not guarantee real requests stay inside it. Real steps vary — some fast, some slow, depending on what else the server is doing right then. A budget is a target and an early-warning system, not a promise that every single request will hit it.

Remember this

  • A latency budget splits the total allowed time across every step of a request.
  • It turns "the service feels slow" into "step X is taking too long" — a much easier problem to fix.
  • Meeting the budget on average is not the same as meeting it for every request. Tail latency covers that gap.

What to learn next

Developer — Code and libraries.

Setup

No installation needed — this uses only Python's built-in time and random modules.

Measuring where a real budget actually goes

latency_budget.py
import random
import time

# Mean time each stage takes, in seconds. Stand-ins for real work -- each
# function actually sleeps for close to this long, with jitter, so the
# measurements below are real elapsed time, not invented numbers.
STAGE_MEANS = {
    "auth_check": 0.003,
    "fetch_features": 0.008,
    "model_call": 0.012,
    "format_response": 0.001,
}
BUDGET = 0.030   # the whole request must finish inside 30ms


def run_stage(mean):
    jittered = max(0.0005, random.gauss(mean, mean * 0.35))
    time.sleep(jittered)
    return jittered


def handle_request():
    timings = {}
    for stage, mean in STAGE_MEANS.items():
        timings[stage] = run_stage(mean)
    return timings


if __name__ == "__main__":
    random.seed(7)
    N = 150
    totals = []
    per_stage = {stage: [] for stage in STAGE_MEANS}
    breaches = 0

    for _ in range(N):
        timings = handle_request()
        total = sum(timings.values())
        totals.append(total)
        for stage, t in timings.items():
            per_stage[stage].append(t)
        if total > BUDGET:
            breaches += 1

    print(f"budget: {BUDGET*1000:.0f}ms   requests: {N}   over budget: {breaches} ({breaches/N*100:.1f}%)")
    print()
    print(f"{'stage':16s} {'avg ms':>8s} {'share of budget':>16s}")
    for stage, values in per_stage.items():
        avg_ms = sum(values) / len(values) * 1000
        share = avg_ms / (BUDGET * 1000) * 100
        print(f"{stage:16s} {avg_ms:8.2f} {share:15.1f}%")
    print(f"{'TOTAL':16s} {sum(totals)/N*1000:8.2f}")
Output
budget: 30ms   requests: 150   over budget: 22 (14.7%)

stage              avg ms  share of budget
auth_check           3.00            10.0%
fetch_features       8.32            27.7%
model_call          12.37            41.2%
format_response      0.98             3.3%
TOTAL               24.67

Real measurements from this exact script — the jitter is generated by random.gauss with a fixed seed, so this exact run is reproducible, but your own machine's raw sleep timings will differ slightly. The finding worth noticing: the average total (24.67ms) comfortably beats the 30ms budget, yet 14.7% of individual requests still went over it. Averages hide exactly the requests you most need to see.

Line-by-line walkthrough

run_stage adds jitter with random.gauss instead of a fixed sleep, because real steps never take exactly the same time twice — this makes the simulation behave more like a real service.

The breaches counter checks each request's own total against BUDGET, separately from the average. That distinction — average versus per-request — is the whole point of this demo.

Common mistakes

Setting a budget from the average alone. A budget built only from average step times will be breached by a real, meaningful share of requests, as shown above. Build it with headroom, and check it against the tail, not the mean.

Forgetting network time between steps. If steps live in different services, the network hop between them belongs in the budget too. It is easy to budget only the code you can see.

Never revisiting the budget as the system changes. A budget written once, at launch, drifts out of date as features get added. Treat it as a living document, checked against real measurements periodically.

Try it yourself

Change model_call's mean from 0.012 to 0.018, simulating a model that got slower after a retrain, and rerun. Watch how much the breach rate rises, even though only one stage changed.

What to learn next

Researcher — Mathematics and papers.

Budgets as a constrained allocation problem

Given a total budget $B$ and $n$ stages with mean latencies $\mu_1, \ldots, \mu_n$, a naive allocation sets each stage's target to $\mu_i$ and hopes $\sum \mu_i \leq B$. This ignores variance entirely, and is why the demo above breaches budget on nearly one request in seven despite the mean total sitting well under it.

A more defensible allocation reserves headroom proportional to each stage's variance, not only its mean — a stage with high variance (a downstream call over an unreliable network) deserves a larger safety margin than a stage that is consistently fast, even if their means are identical.

Composing budgets across independent stages

If stage latencies are independent, the variance of the total is the sum of the stage variances: $\mathrm{Var}(\sum X_i) = \sum \mathrm{Var}(X_i)$. This holds regardless of each stage's individual distribution shape, which is convenient — but the independence assumption itself frequently fails in practice, since shared resource contention (CPU, a connection pool) correlates stage latencies under load, making tails worse than an independence assumption would predict.

Relationship to SLOs

A latency budget is an internal engineering tool for allocating time across a request's steps. A service-level objective is an external commitment about the whole request, typically stated as a percentile target over a time window. The budget is how you engineer toward meeting the SLO; the SLO is how you decide whether you succeeded.

References

  • Dean, J. and Barroso, L., The Tail at Scale, CACM 2013 — the foundational treatment of why per-stage variance, not only per-stage mean, dominates system-level tail behaviour.
  • Google SRE Workbook, Chapter 5, SLO Engineering Case Studies — practical examples of decomposing an end-to-end target into per-component budgets.

What to learn next

What to learn next

These follow on from what you just read.

  • Latency, Load Testing and Capacity

    Tail latency and percentiles

    The average response time can look perfectly healthy while a real slice of your users wait far longer, which is why percentiles like p95 and p99 matter more than the mean.

  • Latency, Load Testing and Capacity

    Load testing a model endpoint

    Load testing sends real, concurrent traffic at a service before real users do, so you find its breaking point on your own schedule instead of during a launch.

  • Latency, Load Testing and Capacity

    Coordinated omission

    A naive load test that waits for each response before sending the next request can completely hide a real outage, because it stops sampling exactly when things are going wrong.