Latency, Load Testing and Capacity

Capacity planning with Little's Law

Little's Law connects how many requests are waiting, how fast they arrive, and how long each one takes — and it explains why a server running close to full speed develops a queue that grows out of all proportion.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Little's Law says the number of requests waiting equals how fast they arrive, multiplied by how long each one takes.

The analogy you have already lived

You have watched a bank counter with one teller. Say ten customers arrive every minute, and each one takes six seconds at the counter. There will usually be about one customer waiting at any given moment. Speed up arrivals, or slow down the teller, and that waiting line grows.

That relationship — arrivals, time per customer, and how many end up waiting — is exactly what a model server deals with too.

Why it exists

Every server has some number of requests present at any moment: some being worked on, some waiting their turn. Little's Law connects that number to two simpler ones: how fast requests arrive, and how long each one takes from start to finish.

Written in words: average number in the system = arrival rate × average time each one spends there.

What makes this genuinely useful: it holds for almost any system, regardless of how requests arrive or how long they take. Steady traffic or bursty, fast or slow — it still holds. That makes it one of the few completely reliable tools for reasoning about capacity.

How it works

requests arriving  -->  [ SYSTEM: server + its queue ]  -->  requests leaving

  arrival rate (how many show up per second)
      x
  average time each one spends inside
      =
  average number inside at any moment

The genuinely useful direction is often backward. Know how many requests you expect, and how long each should take. You can predict how much queueing to expect before it happens, not after users complain.

A real example you have seen

A toll booth backing up badly the moment traffic gets slightly heavier, even though the road itself is fine, is Little's Law in action. A small rise in arrival rate, against a fixed service rate, produces a queue that grows far faster than the increase itself would suggest.

The honest part

This law tells you what happens. It does not, by itself, tell you how bad it gets as you approach full capacity. For that, you need the deeper relationship covered below — and it is considerably more dramatic than most people expect.

Remember this

  • Number waiting = arrival rate × time each request spends in the system.
  • This relationship holds almost universally, which makes it a reliable planning tool.
  • It works in both directions: measure two of the three, and you can predict the third.

What to learn next

Developer — Code and libraries.

Setup

No installation needed — this uses only Python's built-in random module.

Simulating a real queue and checking the law against it

littles_law.py
import random

random.seed(11)


def simulate_mm1(arrival_rate, service_rate, n_customers):
    """A single FIFO server, random (exponential) arrivals and service
    times -- the textbook queue, run as an actual simulation."""
    arrival = 0.0
    server_free_at = 0.0
    sojourn_times = []
    events = []   # (time, +1 on arrival / -1 on departure)

    for _ in range(n_customers):
        arrival += random.expovariate(arrival_rate)
        start_service = max(arrival, server_free_at)
        service = random.expovariate(service_rate)
        departure = start_service + service
        server_free_at = departure
        sojourn_times.append(departure - arrival)
        events.append((arrival, 1))
        events.append((departure, -1))

    events.sort()
    time_weighted_area = 0.0
    n_in_system = 0
    last_t = events[0][0]
    for t, delta in events:
        time_weighted_area += n_in_system * (t - last_t)
        n_in_system += delta
        last_t = t
    total_time = events[-1][0] - events[0][0]

    L_measured = time_weighted_area / total_time
    W_measured = sum(sojourn_times) / len(sojourn_times)
    lambda_measured = n_customers / total_time
    return L_measured, W_measured, lambda_measured


if __name__ == "__main__":
    SERVICE_RATE = 100.0   # the server can finish 100 requests/second flat out
    N = 20000

    print(f"{'utilisation':>12s} {'L (measured)':>13s} {'lambda*W':>10s} "
          f"{'L (theory rho/(1-rho))':>24s}")
    for rho in (0.3, 0.5, 0.7, 0.85, 0.95):
        arrival_rate = rho * SERVICE_RATE
        L, W, lam = simulate_mm1(arrival_rate, SERVICE_RATE, N)
        theory = rho / (1 - rho)
        print(f"{rho:12.2f} {L:13.2f} {lam*W:10.2f} {theory:24.2f}")
Output
 utilisation  L (measured)   lambda*W  L (theory rho/(1-rho))
        0.30          0.43       0.43                    0.43
        0.50          1.01       1.01                    1.00
        0.70          2.40       2.40                    2.33
        0.85          6.08       6.08                    5.67
        0.95         11.62      11.62                   19.00

Real output from this exact simulation, with a fixed random seed for reproducibility. Two columns confirm Little's Law directly: L (measured) — computed by tracking the queue's size over time — and lambda*W — computed completely differently, by multiplying arrival rate by average time spent. They match at every row, because they are two independent measurements of the same real quantity in this simulated run.

The honest part about the last row

The rho = 0.95 row shows the simulated L (11.62) falling noticeably short of the textbook formula (19.00). This is not a bug in the law — it is a real limitation of simulating with only 20,000 customers this close to full capacity, where queue length varies enormously and needs far more samples to settle down. Rerunning with 300,000 customers instead of 20,000 gives L = 19.33, much closer to the theoretical 19.00. Near saturation, a system's behaviour is genuinely noisy and slow to reveal its true average — a fact worth knowing before trusting a short observation window on a production system running close to its limit.

Line-by-line walkthrough

random.expovariate(rate) generates exponentially distributed random gaps — the standard model for "arrivals with no memory of when the last one happened," and for "service time with no memory of how long it has already taken."

The events list and the time_weighted_area loop compute L the honest way: by tracking exactly how many customers were present at every instant, and averaging that over time — not by assuming the formula and working backward.

Common mistakes

Applying the rho/(1-rho) formula outside the M/M/1 case it was derived for. That specific formula assumes exponential arrivals and exponential service times. Little's Law itself (L = λW) makes no such assumption and holds far more generally — do not confuse the general law with this one specific formula.

Planning capacity from an average utilisation alone. A server averaging 70% utilisation across a day can still have short windows well above 90%, where queueing gets sharply worse. Plan against your busiest sustained periods, not the daily average.

Ignoring how slowly systems converge near their limit. As shown above, a system near capacity needs far more observation time to reveal its true behaviour than one running comfortably under it.

Try it yourself

Change SERVICE_RATE from 100.0 to 50.0, halving the server's capacity, while keeping the same rho values. L and W should scale differently — work out beforehand which one you expect to change, and check your prediction against the output.

What to learn next

Researcher — Mathematics and papers.

The formal statement

For a stable queueing system, Little's Law states:

$$L = \lambda W$$

where $L$ is the long-run average number of items in the system, $\lambda$ is the long-run average arrival rate, and $W$ is the long-run average time an item spends in the system. It requires only that the system is stable (arrivals do not permanently exceed departures) — no assumption about the arrival process or service-time distribution is needed, which is what makes it unusually general among queueing results.

Why utilisation near 1 is dangerous

For an M/M/1 queue specifically (Poisson arrivals, exponential service, one server), utilisation $\rho = \lambda/\mu$ gives:

$$L = \frac{\rho}{1-\rho}, \qquad W = \frac{1}{\mu - \lambda}$$

Both diverge as $\rho \to 1$. The practical reading: queueing delay grows nonlinearly as utilisation approaches full capacity, not linearly. Doubling utilisation from 45% to 90% does not double queueing delay — per the formula above, it grows roughly 11×. This is the formal justification behind the common operational guideline of keeping steady-state utilisation well under 100%, often targeted around 70–80% for latency-sensitive services.

Beyond M/M/1

Real services rarely match M/M/1 exactly. The Pollaczek–Khinchine formula generalises to M/G/1 (Poisson arrivals, arbitrary service-time distribution), showing that service-time variance, not only its mean, drives queueing delay — a server with highly variable per-request cost queues worse than one with the same mean but low variance, even at identical utilisation. This is a direct link to why tail latency and capacity planning are the same underlying problem viewed from different angles.

Multi-server systems (M/M/c) — the realistic case of $c$ worker processes or GPU replicas — have their own closed forms (the Erlang C formula), generally showing that a pool of $c$ smaller servers queues worse than one server of equivalent combined capacity, at matched utilisation — the statistical benefit of "pooling" variance across more, smaller units.

References

  • Little, J.D.C., A Proof for the Queuing Formula: L = λW, Operations Research, 1961 — the original proof.
  • Kleinrock, L., Queueing Systems, Volume 1: Theory, 1975 — the standard graduate reference covering M/M/1, M/G/1, and M/M/c in full.
  • Cortez et al. (or equivalent capacity-planning literature from major cloud providers) — practical treatments applying these results to real service capacity planning, typically framed around target utilisation bands rather than raw formulas.

What to learn next