Latency, Load Testing and Capacity

Tail latency and percentiles

The average response time can look perfectly healthy while a real slice of your users wait far longer, which is why percentiles like p95 and p99 matter more than the mean.

Read these first

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

The average response time can look fine while a real slice of your users wait far longer than that average.

The analogy you have already lived

You have waited for a bus that is usually on time. Once in a while it arrives forty minutes late, because of a traffic jam somewhere upstream. If someone asked "how long do you usually wait?", the honest average answer sounds fine. It says nothing about the day you nearly missed a flight because of that one late bus.

Response times work the same way. The average hides the bad days.

Why it exists

The mean, or average, adds up every response time and divides by how many there were. It is easy to compute, and easy to misread. A handful of very slow requests can sit quietly inside a mean that still looks small, diluted by thousands of fast ones.

A percentile answers a sharper question: what response time does a given fraction of requests fall under? The p50 — the median — is the point where half of requests were faster and half were slower. The p99 is the point where 99% of requests were faster, and the slowest 1% were not.

That slowest 1% is the tail. It is small in count, and it is exactly what a frustrated user remembers.

How it works

5000 response times, sorted from fastest to slowest:

fastest ................................................... slowest
|-------------------------------------------------------------|
        ^p50        ^p90    ^p95              ^p99      ^p999
       5ms          6ms     7ms              180ms     240ms

Most requests cluster on the left. A small tail stretches far to the
right -- and that tail is invisible if you only look at the average.

A real example you have seen

A food delivery app's "usually arrives in 30 minutes" can be entirely true on average. But the one time your food took ninety minutes is the experience you remember and complain about. Companies that take this seriously report the slow cases, not only the typical one.

The honest part

There is no single "right" percentile to watch. p50 tells you about the typical user. p99 tells you about the unlucky one in a hundred. Both are real, and they can move in opposite directions — a change that makes the typical case faster can sometimes make the worst case slower. Watch more than one number.

Remember this

  • The mean can look healthy while a real share of users have a bad experience.
  • Percentiles — p50, p95, p99 — describe specific points in the distribution, not an overall blur.
  • The slowest requests are the tail, and they are usually the ones users remember.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Seeing the gap between mean and tail

tail_latency.py
import numpy as np

rng = np.random.default_rng(seed=42)
N = 5000

# A synthetic latency sample, not a real measurement: 98% of requests are
# quick, 2% hit something slow -- a lock, a GC pause, a cold cache row.
fast = rng.normal(loc=5, scale=1.0, size=int(N * 0.98))
slow = rng.normal(loc=180, scale=40.0, size=int(N * 0.02))
latencies_ms = np.clip(np.concatenate([fast, slow]), 0.5, None)
rng.shuffle(latencies_ms)

mean = latencies_ms.mean()
p50, p90, p95, p99, p999 = np.percentile(latencies_ms, [50, 90, 95, 99, 99.9])
worst = latencies_ms.max()

print(f"n={N} synthetic requests")
print(f"mean  : {mean:6.1f} ms")
print(f"p50   : {p50:6.1f} ms")
print(f"p90   : {p90:6.1f} ms")
print(f"p95   : {p95:6.1f} ms")
print(f"p99   : {p99:6.1f} ms")
print(f"p999  : {p999:6.1f} ms")
print(f"worst : {worst:6.1f} ms")
print()
print(f"requests slower than the mean: "
      f"{(latencies_ms > mean).sum()} / {N} "
      f"({(latencies_ms > mean).mean()*100:.1f}%)")
Output
n=5000 synthetic requests
mean  :    8.5 ms
p50   :    5.0 ms
p90   :    6.4 ms
p95   :    6.9 ms
p99   :  182.8 ms
p999  :  239.8 ms
worst :  264.2 ms

requests slower than the mean: 100 / 5000 (2.0%)

This data is synthetically generated with a fixed random seed, not measured against a real service — but the numbers are real outputs of this exact script, and reproducible if you run it again. Notice that mean (8.5ms) sits above p95 (6.9ms). A small, slow tail can drag the average past the 95th percentile itself, which is a genuinely counterintuitive result worth sitting with.

Line-by-line walkthrough

fast and slow model two different populations of requests, mixed together — realistic, because slowness in real services usually comes from a specific rare cause, not a smooth spread.

np.percentile(latencies_ms, [50, 90, 95, 99, 99.9]) computes every percentile in one call, sorting the array once rather than five separate times.

Common mistakes

Reporting only the mean in a dashboard. It is the single most common mistake in this area, and it actively hides the exact problem a percentile view would show.

Averaging percentiles across machines. Averaging five servers' p99 values does not give you the fleet's real p99 — it can be off by a wide margin. Percentiles need to be computed over the combined raw data, or estimated carefully with a merge-aware structure like a t-digest.

Picking p99 with too few samples. With only 50 requests, "p99" is really only the single slowest one, and it is mostly noise. Percentiles need real volume behind them to mean anything.

Confusing p99 with "the worst case." p99 still ignores the worst 1%. If you care about that slice specifically, look at p999 or the true maximum, not p99.

Try it yourself

Change the slow group's size from int(N * 0.02) to int(N * 0.10) — a much more common slow path — and rerun. Watch how far down the percentile list the effect now reaches.

What to learn next

Researcher — Mathematics and papers.

Percentile definitions and interpolation

For a sorted sample of size $n$, the $p$-th percentile has several standard definitions that disagree slightly at small $n$ — linear interpolation between the two nearest ranks (NumPy's default, method="linear"), the nearest-rank method, and others defined in Hyndman and Fan (1996), which catalogues nine variants used across different statistical software. For large $n$ the differences are negligible; for the tiny samples common in per-minute service dashboards, the choice of method can shift a reported p99 measurably.

Approximate percentiles at scale

Computing an exact percentile requires sorting, or at minimum holding the full sample — impractical for a service processing millions of requests per second across a fleet. Two structures dominate in practice:

  • t-digest (Dunning and Ertl, 2019) — a sketch that clusters extreme values more finely than central ones, giving low relative error specifically where percentile queries matter most, and mergeable across shards without recomputing from raw data.
  • HdrHistogram (Tene, 2014) — a fixed-precision histogram over a bounded value range, offering guaranteed relative error and O(1) memory, at the cost of needing a sensible range configured up front.

Both are designed to be mergeable — a property naive per-machine percentile averaging does not have, and the reason "average the p99s" is a genuine measurement error, not a shortcut.

The tail's effect on composed systems

For a request that fans out to $k$ independent backend calls and waits for all of them, the probability that at least one lands beyond the individual p99 is $1 - 0.99^{k}$ — about 63% once $k = 100$. This is the core argument of Dean and Barroso (2013): tail latency compounds with fan-out, so a service built from many small calls has a worse effective tail than any single call in it, even when every individual call has an excellent p99.

References

  • Dean, J. and Barroso, L., The Tail at Scale, CACM 2013 — research.google/pubs/the-tail-at-scale
  • Hyndman, R. and Fan, Y., Sample Quantiles in Statistical Packages, The American Statistician, 1996
  • Dunning, T. and Ertl, O., Computing Extremely Accurate Quantiles Using t-Digests, 2019 — arxiv.org/abs/1902.04023
  • Tene, G., HdrHistogram: A High Dynamic Range Histogram, 2014 — hdrhistogram.org

What to learn next