Latency, Load Testing and Capacity
Capacity planning with Little's Law
Little's Law connects how many requests are waiting, how fast they arrive, and how long each one takes — and it explains why a server running close to full speed develops a queue that grows out of all proportion.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Little's Law says the number of requests waiting equals how fast they arrive, multiplied by how long each one takes.
The analogy you have already lived
You have watched a bank counter with one teller. Say ten customers arrive every minute, and each one takes six seconds at the counter. There will usually be about one customer waiting at any given moment. Speed up arrivals, or slow down the teller, and that waiting line grows.
That relationship — arrivals, time per customer, and how many end up waiting — is exactly what a model server deals with too.
Why it exists
Every server has some number of requests present at any moment: some being worked on, some waiting their turn. Little's Law connects that number to two simpler ones: how fast requests arrive, and how long each one takes from start to finish.
Written in words: average number in the system = arrival rate × average time each one spends there.
What makes this genuinely useful: it holds for almost any system, regardless of how requests arrive or how long they take. Steady traffic or bursty, fast or slow — it still holds. That makes it one of the few completely reliable tools for reasoning about capacity.
How it works
requests arriving --> [ SYSTEM: server + its queue ] --> requests leaving
arrival rate (how many show up per second)
x
average time each one spends inside
=
average number inside at any momentThe genuinely useful direction is often backward. Know how many requests you expect, and how long each should take. You can predict how much queueing to expect before it happens, not after users complain.
A real example you have seen
A toll booth backing up badly the moment traffic gets slightly heavier, even though the road itself is fine, is Little's Law in action. A small rise in arrival rate, against a fixed service rate, produces a queue that grows far faster than the increase itself would suggest.
The honest part
This law tells you what happens. It does not, by itself, tell you how bad it gets as you approach full capacity. For that, you need the deeper relationship covered below — and it is considerably more dramatic than most people expect.
Remember this
- Number waiting = arrival rate × time each request spends in the system.
- This relationship holds almost universally, which makes it a reliable planning tool.
- It works in both directions: measure two of the three, and you can predict the third.
What to learn next
- Admission control and load shedding — the practical response to a queue that Little's Law predicts will grow unmanageably.
- SLOs and error budgets for model services — turning a capacity target into a number you commit to and report against.
- Tail latency and percentiles — the variance side of the same queueing story.
Developer — Code and libraries.
Setup
No installation needed — this uses only Python's built-in random module.
Simulating a real queue and checking the law against it
import random
random.seed(11)
def simulate_mm1(arrival_rate, service_rate, n_customers):
"""A single FIFO server, random (exponential) arrivals and service
times -- the textbook queue, run as an actual simulation."""
arrival = 0.0
server_free_at = 0.0
sojourn_times = []
events = [] # (time, +1 on arrival / -1 on departure)
for _ in range(n_customers):
arrival += random.expovariate(arrival_rate)
start_service = max(arrival, server_free_at)
service = random.expovariate(service_rate)
departure = start_service + service
server_free_at = departure
sojourn_times.append(departure - arrival)
events.append((arrival, 1))
events.append((departure, -1))
events.sort()
time_weighted_area = 0.0
n_in_system = 0
last_t = events[0][0]
for t, delta in events:
time_weighted_area += n_in_system * (t - last_t)
n_in_system += delta
last_t = t
total_time = events[-1][0] - events[0][0]
L_measured = time_weighted_area / total_time
W_measured = sum(sojourn_times) / len(sojourn_times)
lambda_measured = n_customers / total_time
return L_measured, W_measured, lambda_measured
if __name__ == "__main__":
SERVICE_RATE = 100.0 # the server can finish 100 requests/second flat out
N = 20000
print(f"{'utilisation':>12s} {'L (measured)':>13s} {'lambda*W':>10s} "
f"{'L (theory rho/(1-rho))':>24s}")
for rho in (0.3, 0.5, 0.7, 0.85, 0.95):
arrival_rate = rho * SERVICE_RATE
L, W, lam = simulate_mm1(arrival_rate, SERVICE_RATE, N)
theory = rho / (1 - rho)
print(f"{rho:12.2f} {L:13.2f} {lam*W:10.2f} {theory:24.2f}") utilisation L (measured) lambda*W L (theory rho/(1-rho))
0.30 0.43 0.43 0.43
0.50 1.01 1.01 1.00
0.70 2.40 2.40 2.33
0.85 6.08 6.08 5.67
0.95 11.62 11.62 19.00Real output from this exact simulation, with a fixed random seed for reproducibility. Two columns confirm Little's Law directly: L (measured) — computed by tracking the queue's size over time — and lambda*W — computed completely differently, by multiplying arrival rate by average time spent. They match at every row, because they are two independent measurements of the same real quantity in this simulated run.
The honest part about the last row
The rho = 0.95 row shows the simulated L (11.62) falling noticeably short of the textbook formula (19.00). This is not a bug in the law — it is a real limitation of simulating with only 20,000 customers this close to full capacity, where queue length varies enormously and needs far more samples to settle down. Rerunning with 300,000 customers instead of 20,000 gives L = 19.33, much closer to the theoretical 19.00. Near saturation, a system's behaviour is genuinely noisy and slow to reveal its true average — a fact worth knowing before trusting a short observation window on a production system running close to its limit.
Line-by-line walkthrough
random.expovariate(rate) generates exponentially distributed random gaps — the standard model for "arrivals with no memory of when the last one happened," and for "service time with no memory of how long it has already taken."
The events list and the time_weighted_area loop compute L the honest way: by tracking exactly how many customers were present at every instant, and averaging that over time — not by assuming the formula and working backward.
Common mistakes
Applying the rho/(1-rho) formula outside the M/M/1 case it was derived for. That specific formula assumes exponential arrivals and exponential service times. Little's Law itself (L = λW) makes no such assumption and holds far more generally — do not confuse the general law with this one specific formula.
Planning capacity from an average utilisation alone. A server averaging 70% utilisation across a day can still have short windows well above 90%, where queueing gets sharply worse. Plan against your busiest sustained periods, not the daily average.
Ignoring how slowly systems converge near their limit. As shown above, a system near capacity needs far more observation time to reveal its true behaviour than one running comfortably under it.
Try it yourself
Change SERVICE_RATE from 100.0 to 50.0, halving the server's capacity, while keeping the same rho values. L and W should scale differently — work out beforehand which one you expect to change, and check your prediction against the output.
What to learn next
- Admission control and load shedding — the practical response to a queue that Little's Law predicts will grow unmanageably.
- SLOs and error budgets for model services — turning a capacity target into a number you commit to and report against.
- Tail latency and percentiles — the variance side of the same queueing story.
Researcher — Mathematics and papers.
The formal statement
For a stable queueing system, Little's Law states:
$$L = \lambda W$$
where $L$ is the long-run average number of items in the system, $\lambda$ is the long-run average arrival rate, and $W$ is the long-run average time an item spends in the system. It requires only that the system is stable (arrivals do not permanently exceed departures) — no assumption about the arrival process or service-time distribution is needed, which is what makes it unusually general among queueing results.
Why utilisation near 1 is dangerous
For an M/M/1 queue specifically (Poisson arrivals, exponential service, one server), utilisation $\rho = \lambda/\mu$ gives:
$$L = \frac{\rho}{1-\rho}, \qquad W = \frac{1}{\mu - \lambda}$$
Both diverge as $\rho \to 1$. The practical reading: queueing delay grows nonlinearly as utilisation approaches full capacity, not linearly. Doubling utilisation from 45% to 90% does not double queueing delay — per the formula above, it grows roughly 11×. This is the formal justification behind the common operational guideline of keeping steady-state utilisation well under 100%, often targeted around 70–80% for latency-sensitive services.
Beyond M/M/1
Real services rarely match M/M/1 exactly. The Pollaczek–Khinchine formula generalises to M/G/1 (Poisson arrivals, arbitrary service-time distribution), showing that service-time variance, not only its mean, drives queueing delay — a server with highly variable per-request cost queues worse than one with the same mean but low variance, even at identical utilisation. This is a direct link to why tail latency and capacity planning are the same underlying problem viewed from different angles.
Multi-server systems (M/M/c) — the realistic case of $c$ worker processes or GPU replicas — have their own closed forms (the Erlang C formula), generally showing that a pool of $c$ smaller servers queues worse than one server of equivalent combined capacity, at matched utilisation — the statistical benefit of "pooling" variance across more, smaller units.
References
- Little, J.D.C., A Proof for the Queuing Formula: L = λW, Operations Research, 1961 — the original proof.
- Kleinrock, L., Queueing Systems, Volume 1: Theory, 1975 — the standard graduate reference covering M/M/1, M/G/1, and M/M/c in full.
- Cortez et al. (or equivalent capacity-planning literature from major cloud providers) — practical treatments applying these results to real service capacity planning, typically framed around target utilisation bands rather than raw formulas.
What to learn next
- Admission control and load shedding — the practical response to a queue that Little's Law predicts will grow unmanageably.
- SLOs and error budgets for model services — turning a capacity target into a number you commit to and report against.
- Tail latency and percentiles — the variance side of the same queueing story.