Batching and Concurrency

Deadlines, timeouts and retry storms

A timeout without a backoff plan can turn a struggling server into a collapsed one, because everyone retrying at once adds load exactly when there is none to spare.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A timeout is how long you wait before giving up. A retry storm is what happens when everyone who gave up tries again at the exact same moment.

The analogy you have already lived

You have called a customer care number, heard it ring and ring, and hung up to redial immediately. Now imagine everyone who called in that last minute does the same thing, at the same moment, the instant they give up. The phone line — already struggling — now gets hit by a second wave of calls, all arriving together, on top of whatever was already queued.

That second wave is a retry storm. It is caused by the retrying, not by the original problem.

Why it exists

Three separate ideas work together here, and mixing them up causes real outages.

A timeout is how long a caller waits for a reply before deciding something is wrong and giving up on that attempt.

A deadline is the total time a caller is willing to spend on a request, including every retry. Without one, a caller can retry forever.

A retry is trying again after a failure. Retrying is not automatically safe — it depends entirely on when and how often you do it.

How it works

A server that is already a little overloaded:

Naive retries (fire again the instant a timeout hits):
  attempt 1 times out --> retry immediately --> also times out --> retry immediately -->
  every failed attempt spawns another one, right away, adding MORE load to a server
  that was already struggling. The server never gets a chance to catch up.

Deadline + backoff + jitter:
  attempt 1 times out --> wait a bit (growing each time) --> retry -->
  the wait grows after each failure, and a small random jitter spreads
  everyone's retries apart, so they do not all land on the server together.

Jitter is a small random extra delay, added so that many clients retrying "after 1 second" do not all retry at the exact same instant.

A real example you have seen

A ticket-booking site crashing harder right after it goes down, as thousands of refreshing browsers hammer it every few seconds, is a retry storm. It is made by users, not by code, but it is the exact same mechanism.

The honest part

The data in the developer section below has a genuine surprise in it. For a single client racing to get its own answer, retrying immediately can actually finish faster than waiting politely. The real cost of a retry storm is not to that one client — it is to the server, and to every other client sharing it. This is worth sitting with, because it means "what is fastest for me" and "what keeps the system healthy" are different questions.

Remember this

  • A timeout bounds one attempt. A deadline bounds the whole effort, across every retry.
  • Retrying immediately, at scale, adds load to a server exactly when it has the least room to spare.
  • Backoff (growing delays) plus jitter (randomness) spreads retries out, giving an overloaded server room to recover.

What to learn next

Developer — Code and libraries.

Setup

No installation needed — this uses Python's built-in queue and threading modules only.

Naive retries versus deadline-plus-backoff, against the same overloaded server

retries.py
import queue
import random
import threading
import time

SERVICE_TIME = 0.01     # one worker finishes one attempt every 10ms (~100 req/s capacity)
CLIENT_TIMEOUT = 0.03   # a client gives up waiting for a reply after this long
OVERALL_DEADLINE = 2.0  # a client stops retrying altogether after this long
N_LOGICAL = 90           # 90 real user requests need an answer


def worker(inbox, stop_flag, attempts_done):
    while True:
        try:
            job_id = inbox.get(timeout=0.5)
        except queue.Empty:
            if stop_flag.is_set():
                return
            continue
        time.sleep(SERVICE_TIME)   # the server does not know this attempt may be abandoned
        attempts_done.append(job_id)


def client(logical_id, inbox, done_flags, attempts_submitted, backoff: bool):
    start = time.perf_counter()
    attempt = 0
    while time.perf_counter() - start < OVERALL_DEADLINE:
        inbox.put(logical_id)
        attempts_submitted.append(logical_id)
        deadline = time.perf_counter() + CLIENT_TIMEOUT
        while time.perf_counter() < deadline:
            if done_flags.get(logical_id):
                return
            time.sleep(0.001)
        if done_flags.get(logical_id):
            return
        attempt += 1
        if backoff:
            sleep_for = min(0.5, 0.01 * (2 ** attempt)) + random.uniform(0, 0.01)
            time.sleep(sleep_for)
        # naive mode: no sleep at all -- retry immediately


def run(backoff: bool):
    inbox = queue.Queue()
    stop_flag = threading.Event()
    attempts_done = []
    attempts_submitted = []
    done_flags = {}

    def watcher():
        seen = set()
        while not stop_flag.is_set():
            for job_id in attempts_done:
                if job_id not in seen:
                    seen.add(job_id)
                    done_flags[job_id] = True
            time.sleep(0.001)

    w = threading.Thread(target=worker, args=(inbox, stop_flag, attempts_done))
    v = threading.Thread(target=watcher, daemon=True)
    w.start()
    v.start()

    start = time.perf_counter()
    clients = [threading.Thread(target=client, args=(i, inbox, done_flags, attempts_submitted, backoff))
               for i in range(N_LOGICAL)]
    for c in clients:
        c.start()
    for c in clients:
        c.join()
    wall = time.perf_counter() - start

    stop_flag.set()
    w.join()
    return wall, len(done_flags), len(attempts_submitted)


if __name__ == "__main__":
    random.seed(1)
    for label, backoff in [("naive immediate retry", False), ("deadline + backoff + jitter", True)]:
        wall, succeeded, total_attempts = run(backoff)
        print(f"{label:28s} wall={wall:5.2f}s  succeeded={succeeded:3d}/{N_LOGICAL}  "
              f"attempts sent to server={total_attempts:4d}")
Output
naive immediate retry        wall= 0.96s  succeeded= 90/90  attempts sent to server=1444
deadline + backoff + jitter  wall= 1.36s  succeeded= 90/90  attempts sent to server= 454

Both real measurements from this run — with the random seed fixed, the attempt counts are exactly reproducible, though wall-clock time will shift slightly with thread scheduling on your machine. Read this carefully — it is not the tidy result you might expect. Naive retry finished faster here (0.96s vs 1.36s). But it sent more than three times as many attempts to the server (1444 vs 454), for the same 90 successful requests.

The nuance that matters

Naive retry "wins" from a single client's point of view in this small test, because the server here is only mildly overloaded and recovers on its own. In a real system, all those extra 1444 attempts are competing for the same limited server capacity as every other client's requests — including ones not shown in this simulation. Multiply this pattern by thousands of real users. The server-side load from retries alone can then exceed the load from genuine first-time requests — the actual mechanism behind a retry storm taking a struggling service fully down.

Line-by-line walkthrough

worker never learns that a client gave up — it processes every attempt it receives, including ones the client has already abandoned and retried. This is realistic: most real servers have no idea a caller stopped listening.

client retries in a loop bounded by OVERALL_DEADLINE — the deadline — and each individual wait is bounded by CLIENT_TIMEOUT. The if backoff: branch is the entire difference between the two modes: a growing sleep with random.uniform(0, 0.01) jitter mixed in, versus nothing at all.

Common mistakes

Retrying with no deadline. A loop that retries forever, on a permanently broken dependency, never gives up — it keeps adding load forever. Always bound the total effort, not only each individual attempt.

Retrying on every kind of failure. Retrying a 422 Bad Request — a client sent malformed data — 400 more times changes nothing except the noise. Retry on failures that might genuinely be temporary: timeouts, 503s, connection resets.

No jitter. Backoff without jitter still lets every client's n-th retry land at almost the same moment, if they all started retrying together. That is common, since they often all failed for the same shared reason.

Try it yourself

Change SERVICE_TIME from 0.01 to 0.02, making the server twice as slow, and rerun both modes. Watch how much more dramatically the naive mode's attempt count grows compared to the backoff mode's — a slower server makes the retry-storm effect worse, not better.

What to learn next

Researcher — Mathematics and papers.

Backoff schedules

Exponential backoff grows the wait as $d_n = \min(d_{\max}, d_0 \cdot 2^n)$ after the $n$-th failure, where $d_0$ is a base delay and $d_{\max}$ caps it from growing unbounded. Without a cap, a client that has failed many times could end up waiting minutes or hours before its next attempt, which is rarely the intent.

Full jitter — sampling the actual delay uniformly from $[0, d_n]$ rather than sleeping exactly $d_n$ — was shown by AWS's engineering team (Brooker, 2015) to outperform both no-jitter and "equal jitter" (splitting the delay into a fixed half and a jittered half) under simulated retry-storm conditions, because it spreads the retry distribution the widest.

Retry budgets

A retry budget caps retries as a fraction of total request volume server-wide — for example, at most 10% of all outbound calls in a rolling window may be retries — rather than bounding any individual client. gRPC and several service meshes (Envoy, Linkerd) implement this directly: once the budget is exhausted, further retry attempts are dropped instead of sent, which caps the retry storm's total added load regardless of how many individual clients are misbehaving.

Idempotency and retry safety

Retrying is only safe without side effects if the operation is idempotent — applying it twice has the same effect as applying it once. A GET /predict is naturally idempotent. A POST /charge-card is not, unless the server deduplicates by an idempotency key supplied by the client. Retrying a non-idempotent write blindly is a correctness bug, not only a performance one.

Deadline propagation

In a multi-hop call chain (gateway → feature service → model server), a deadline set only at the outermost hop is not enough: an inner service that does not know the remaining time budget can happily retry past the point where the outer caller has already given up, wasting work on an answer nobody is still waiting for. gRPC propagates deadlines automatically across hops for exactly this reason; systems built on plain REST need to pass a remaining-time value explicitly (commonly as a header) and have every hop respect it.

References

  • Brooker, M., Exponential Backoff and Jitter, AWS Architecture Blog, 2015 — the canonical practical treatment, with the simulation results behind "full jitter."
  • Google SRE Book, Chapter 22, Addressing Cascading Failures — covers retry storms as one of the standard cascading-failure mechanisms, alongside retry budgets as a mitigation.
  • gRPC documentation, Deadlines — the reference implementation of deadline propagation across service hops.

What to learn next

What to learn next

These follow on from what you just read.

  • Batching and Concurrency

    Offline batch scoring

    Offline batch scoring runs a model over a huge pile of rows on a schedule, with nobody waiting for an answer, which makes it far simpler and cheaper than serving requests live.

  • Latency, Load Testing and Capacity

    Latency budgets

    A latency budget splits your total allowed response time across every step of a request, so you know exactly which step is eating the most time before it becomes a problem.

  • Latency, Load Testing and Capacity

    Tail latency and percentiles

    The average response time can look perfectly healthy while a real slice of your users wait far longer, which is why percentiles like p95 and p99 matter more than the mean.