Latency, Load Testing and Capacity

Coordinated omission

A naive load test that waits for each response before sending the next request can completely hide a real outage, because it stops sampling exactly when things are going wrong.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Coordinated omission is when a load test stops measuring exactly during the moments it should be measuring the most.

The analogy you have already lived

You have stood in a slow government-office queue where the clerk vanishes for a long tea break. The line behind that counter grows and grows. Now imagine someone measuring "wait time" only by asking each person, right as they finally reach the counter, how long they personally waited.

During the tea break, nobody reaches the counter. Nobody gets asked. The survey shows nothing unusual happening — right as the queue outside is at its worst.

Why it exists

A simple way to load-test a service is: send a request, wait for the answer, then send the next one. This is a closed-loop load generator — it waits for each response before sending the next request.

That sounds harmless. It has a serious flaw. If the server stalls for one second, a closed-loop client waits. It does not send nine other requests during that second, the way nine real independent users would. When the server finally responds, the client records one slow measurement, and moves on — as if only one request was ever affected.

In reality, everyone who would have arrived during that stall was also waiting, backed up behind it. A closed-loop test never sees them, because it never sent them. This blind spot is coordinated omission — the load generator's own pacing is quietly leaving out exactly the samples that would show the problem.

How it works

Real, independent users (what actually happens in production):
  a new user arrives every 10ms, no matter what the server is doing
  |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
       during a 500ms stall, 50 of them pile up, all waiting

A closed-loop load tester (naive):
  send request --> wait for reply --> send next request --> wait...
       during that same 500ms stall, it has sent exactly ONE request,
       and records exactly ONE slow measurement for the whole event

A real example you have seen

A ticket-booking site that "looked fine in testing" and then fell over on sale day often failed exactly this way. The test tool politely waited for each response before asking again. Real customers, all arriving at once, did not wait politely for each other.

The honest part

This is one of the more surprising ideas in this whole area, and it fools experienced engineers, not only beginners. A load test can report excellent percentiles while the underlying service was, for real users, badly broken. Read the numbers in the developer section slowly — they are correct, and they still feel wrong at first.

Remember this

  • A closed-loop load test waits for each reply before sending the next request.
  • During a slow patch, it sends fewer requests, and its measurements quietly miss the pile-up.
  • The fix is measuring against when a request should have been sent, not against when the tool got around to sending it.

What to learn next

Developer — Code and libraries.

Setup

No installation needed — this is a deterministic, virtual-clock simulation in plain Python. No real sleeping, no real network — it exists to make the mechanism visible, precisely and reproducibly.

The same stall, measured two different ways

coordinated_omission.py
TARGET_RATE = 100          # 100 scheduled requests per second
DURATION = 2.0              # 2 simulated seconds
SERVICE_TIME = 0.007         # a normal request takes 7ms on the server
STALL_AT_REQUEST = 90        # the server freezes right as this request arrives
STALL_LENGTH = 0.5            # a 500ms GC-style pause


def server_finish_times(arrival_times):
    """A single FIFO worker. Every arrival waits for the worker to be
    free, then takes SERVICE_TIME -- except one request that lands on
    a stall and pays extra."""
    finish = []
    worker_free_at = 0.0
    for i, arrival in enumerate(arrival_times):
        start = max(arrival, worker_free_at)
        service = SERVICE_TIME + (STALL_LENGTH if i == STALL_AT_REQUEST else 0.0)
        worker_free_at = start + service
        finish.append(worker_free_at)
    return finish


def percentile(values, p):
    s = sorted(values)
    return s[min(int(len(s) * p), len(s) - 1)]


if __name__ == "__main__":
    gap = 1.0 / TARGET_RATE
    scheduled = [i * gap for i in range(int(DURATION / gap))]

    # Open-loop / corrected: every intended arrival is honoured, so
    # every user who would show up during the stall is counted, with
    # wait measured against when they SHOULD have been served.
    finish = server_finish_times(scheduled)
    corrected_wait = [f - a for f, a in zip(finish, scheduled)]

    # Closed-loop / naive: the load generator only sends the next
    # request once the previous one has returned. It never builds a
    # backlog, so it never sees what a backlog does to real users.
    naive_wait = []
    clock = 0.0
    while clock < DURATION:
        start = clock
        stalled = abs(start - scheduled[STALL_AT_REQUEST]) < gap
        finish_t = start + SERVICE_TIME + (STALL_LENGTH if stalled else 0.0)
        naive_wait.append(finish_t - start)
        clock = finish_t

    print(f"open-loop  (corrected): {len(corrected_wait):4d} requests   "
          f"p50={percentile(corrected_wait,0.50)*1000:7.1f}ms  "
          f"p95={percentile(corrected_wait,0.95)*1000:7.1f}ms  "
          f"p99={percentile(corrected_wait,0.99)*1000:7.1f}ms")
    print(f"closed-loop (naive)  : {len(naive_wait):4d} requests   "
          f"p50={percentile(naive_wait,0.50)*1000:7.1f}ms  "
          f"p95={percentile(naive_wait,0.95)*1000:7.1f}ms  "
          f"p99={percentile(naive_wait,0.99)*1000:7.1f}ms")

    over_corrected = sum(1 for w in corrected_wait if w > 0.1)
    over_naive = sum(1 for w in naive_wait if w > 0.1)
    print(f"\nrequests waiting over 100ms -- corrected: {over_corrected}   naive: {over_naive}")
Output
open-loop  (corrected):  200 requests   p50=  210.0ms  p95=  480.0ms  p99=  504.0ms
closed-loop (naive)  :  215 requests   p50=    7.0ms  p95=    7.0ms  p99=    7.0ms

requests waiting over 100ms -- corrected: 110   naive: 1

This entire output is deterministic — there is no randomness in this script, so rerunning it gives these exact numbers every time. Read the two rows carefully. The exact same server, with the exact same single stall, produces a p99 of 504ms in the corrected measurement and a p99 of 7.0ms — completely normal — in the naive one. 110 requests genuinely waited over 100ms. The naive method saw 1.

Line-by-line walkthrough

server_finish_times computes what actually happens on the server: a FIFO queue where each arrival waits for the worker, and one unlucky arrival triggers a stall that delays everyone queued behind it too — because worker_free_at carries forward.

corrected_wait measures every scheduled arrival against when it was scheduled to arrive, scheduled[i], not against any notion of "when the client happened to ask." This is the entire fix, in one line: f - a where a is the intended arrival time.

naive_wait recomputes what a closed-loop client would see: its own clock only advances by sending a request and waiting for the reply, so during the 500ms stall it is not sending anything else — and does not know 49 other virtual users would have piled up behind it.

Common mistakes

Trusting a load test tool without checking its arrival model. Not every popular tool is open-loop by default. Check the documentation for whether requests are scheduled against wall-clock time or against the previous response.

Assuming a high request count means good coverage. The naive run above sent more total requests (215 vs 200) than the corrected one, and told a far less accurate story. Volume is not the same as validity.

Reading "excellent p99" as proof of a healthy system. If the measurement method itself has this blind spot, an excellent reported p99 can coexist with a genuinely bad user experience. Check how the number was produced before trusting it.

Try it yourself

Change STALL_AT_REQUEST from a single request to a whole range — stall every request between index 90 and 95 — and rerun. Watch the naive method's numbers barely move, while the corrected ones get dramatically worse.

What to learn next

Researcher — Mathematics and papers.

The formal correction

Given an intended arrival schedule $t_1, t_2, \ldots$ and observed completion times $f_1, f_2, \ldots$, the coordinated-omission-corrected latency for request $i$ is $f_i - t_i$ — response time measured against intended send time, not actual send time. A closed-loop generator that only sends request $i+1$ after $f_i$ conflates "time spent waiting to even be sent" with "server think time," systematically undercounting the former whenever the server is behind schedule.

Gil Tene, who coined the term while building HdrHistogram, additionally proposed a retrospective correction: given a known intended interval $\delta$ between requests, synthetic samples can be added after the fact for every interval a real request "should" have arrived during a stall, each attributed the appropriate partial wait. This lets an already-collected closed-loop dataset be corrected approximately, without rerunning the test — useful when only closed-loop data is available, though inferior to generating open-loop data in the first place.

Why this specifically distorts percentiles, not means

A single missed request during a stall has a bounded effect on a mean — it is one data point among many, diluted like any outlier. Its effect on a percentile, especially the tail percentiles this material cares about most, is far larger: the missing samples are disproportionately the slow ones, which is precisely the region a p99 or p999 is trying to characterise. This is why coordinated omission is discussed specifically in the context of tail latency measurement, rather than as a general statistics caveat.

Tools with correction built in

wrk2 (a fork of wrk) and HdrHistogram-based tooling schedule requests against a fixed target rate and record against intended timestamps by construction, rather than requiring a post-hoc correction. k6 and Gatling support open-loop arrival-rate executors specifically to sidestep this class of error, distinct from their default closed-loop, concurrency-based modes.

References

  • Tene, G., How NOT to Measure Latency, Strange Loop 2015 — the original, widely cited talk introducing the term and the correction technique.
  • Tene, G., HdrHistogram: A High Dynamic Range Histogram, 2014 — hdrhistogram.org — the reference implementation most correction tooling builds on.

What to learn next