Batching and Concurrency

Dynamic batching

Dynamic batching holds a few incoming requests for a short moment so one model call can serve all of them together, trading a little latency for a lot more throughput.

Read these first

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Dynamic batching means waiting a few milliseconds for more requests to arrive, then answering all of them with one model call.

The analogy you have already lived

You have waited for a shared autorickshaw or a shuttle van to fill up. The driver does not leave the moment one passenger sits down. They wait a short, fixed time for two or three more people, then drive everyone at once.

Every passenger reaches their stop a little later than if the van had left instantly. But the driver made the trip once instead of four times, using far less fuel per person.

A model server can do the same thing with requests instead of passengers.

Why it exists

An accelerator — a GPU, or even a CPU running a big model — has a fixed cost for every single call. Loading the input, moving data into place, starting the computation. That fixed cost does not grow much whether you send it one row or fifty rows at once.

Answer one request at a time, and you pay that fixed cost fifty times. Group fifty requests into one call, and you pay it once. The hardware finishes real work in roughly the same total time either way — the fixed cost is what you save.

Batching here means grouping several inputs into a single call. Dynamic means the group is not a fixed size decided in advance. It is built on the fly, from whatever requests happen to show up in a short window.

How it works

requests arrive, one by one:
   r1 ---->  [ waiting room ]
   r2 ---->  [ waiting room ]         after max_wait_ms OR
   r3 ---->  [ waiting room ]  ---->  max_batch_size reached
                                              |
                                              v
                                      one model call, batch of 3
                                              |
                                              v
                                   r1, r2, r3 all get answers back

Two knobs control the wait. A maximum batch size stops the wait once you have enough requests. A maximum wait time stops it once too much time has passed, even with a small batch. Whichever limit is hit first wins.

A real example you have seen

A UPI payment app checking many transactions for fraud during a busy sale. A search box that scores every suggestion as you type, and gets a burst of requests the instant a popular query starts trending. Both are cases with many small, similar requests arriving close together — exactly what dynamic batching is built for.

The honest part

Batching always makes the average request wait a little longer, because it is, by definition, waiting for company. If your service already answers every request in two milliseconds, batching can make things worse for the fastest requests. It only pays off once the accelerator's fixed cost is a meaningful slice of the total time.

Remember this

  • Dynamic batching groups several requests into one model call, using a short wait instead of a fixed schedule.
  • It raises throughput — requests finished per second — at the cost of a little extra latency per request.
  • It matters most when a call's fixed cost is large next to the time spent per item, which is exactly the case for most accelerators.

What to learn next

  • Continuous batching for LLMs — the version of this idea built for text generation, where requests do not all finish at the same time.
  • Model serving — the service this batcher would sit inside.
  • Latency budgets — deciding how much of your time budget batching is allowed to spend.

Developer — Code and libraries.

Setup

bash
pip install fastapi "uvicorn[standard]"

No server is needed to see the effect — the demo below runs entirely on its own, timing a real (if artificial) "model" function.

A dynamic batcher, in under 60 lines

dynamic_batching.py
import queue
import threading
import time

FIXED_OVERHEAD = 0.004   # seconds paid once per call, no matter the batch size
PER_ITEM = 0.0006        # seconds paid per item inside the call


def run_model(batch_size: int) -> float:
    """Stands in for a real model call: it sleeps for the time a real
    call would take, and does no real computation."""
    start = time.perf_counter()
    time.sleep(FIXED_OVERHEAD + PER_ITEM * batch_size)
    return time.perf_counter() - start


def one_at_a_time(n_requests):
    latencies = []
    start = time.perf_counter()
    for _ in range(n_requests):
        t0 = time.perf_counter()
        run_model(1)
        latencies.append(time.perf_counter() - t0)
    return time.perf_counter() - start, latencies


class DynamicBatcher:
    """Collects requests for up to max_wait seconds, or until max_batch
    requests are queued -- whichever happens first."""

    def __init__(self, max_batch=16, max_wait=0.01):
        self.max_batch = max_batch
        self.max_wait = max_wait
        self.inbox = queue.Queue()
        self.latencies = []

    def submit(self):
        done = threading.Event()
        self.inbox.put((time.perf_counter(), done))
        done.wait()

    def serve_forever(self, total_requests):
        served = 0
        while served < total_requests:
            batch = [self.inbox.get()]
            deadline = time.perf_counter() + self.max_wait
            while len(batch) < self.max_batch:
                remaining = deadline - time.perf_counter()
                if remaining <= 0:
                    break
                try:
                    batch.append(self.inbox.get(timeout=remaining))
                except queue.Empty:
                    break
            run_model(len(batch))
            now = time.perf_counter()
            for arrival, done in batch:
                self.latencies.append(now - arrival)
                done.set()
            served += len(batch)


def with_dynamic_batching(n_requests, arrival_gap):
    batcher = DynamicBatcher(max_batch=16, max_wait=0.01)
    server = threading.Thread(target=batcher.serve_forever, args=(n_requests,))
    server.start()

    start = time.perf_counter()
    clients = []
    for _ in range(n_requests):
        t = threading.Thread(target=batcher.submit)
        t.start()
        clients.append(t)
        time.sleep(arrival_gap)
    for t in clients:
        t.join()
    server.join()
    return time.perf_counter() - start, batcher.latencies


if __name__ == "__main__":
    N = 300
    ARRIVAL_GAP = 0.0015   # a new request roughly every 1.5ms

    def summarize(name, total_time, latencies):
        avg_ms = sum(latencies) / len(latencies) * 1000
        throughput = len(latencies) / total_time
        print(f"{name:16s} wall={total_time:6.3f}s  "
              f"throughput={throughput:7.1f} req/s  avg latency={avg_ms:6.2f} ms")

    total_naive, lat_naive = one_at_a_time(N)
    summarize("one at a time", total_naive, lat_naive)

    total_batched, lat_batched = with_dynamic_batching(N, ARRIVAL_GAP)
    summarize("dynamic batch", total_batched, lat_batched)
Output
one at a time    wall= 1.679s  throughput=  178.6 req/s  avg latency=  5.60 ms
dynamic batch    wall= 0.849s  throughput=  353.5 req/s  avg latency= 19.20 ms

These are real measurements from this exact script, not invented figures. They will differ on your machine — thread scheduling and CPU speed both move the numbers — but the shape holds: batching roughly doubled throughput here while roughly tripling average latency. That trade is the whole lesson.

Line-by-line walkthrough

FIXED_OVERHEAD and PER_ITEM model a real accelerator's cost curve: a flat setup cost, plus a small additional cost per item in the batch.

DynamicBatcher.serve_forever is the core idea. It grabs one request, then keeps grabbing more with a shrinking timeout — deadline - time.perf_counter() — until either the batch is full or the deadline passes. queue.Queue.get(timeout=...) handles both cases: it returns a request if one arrives in time, or raises queue.Empty if the deadline passes first.

Each request carries its own threading.Event. The batcher sets it once that request's answer is ready, which wakes up exactly that one waiting client thread — not the others.

Common mistakes

No maximum wait. Waiting only for max_batch to fill means a quiet period leaves early requests stuck indefinitely. Always cap the wait time too.

Batching requests of very different sizes together. A batch mixing a 10-token input with a 10,000-token one wastes most of the batch's capacity on padding. Group similar-sized work when you can.

Testing only under light load. A batcher's whole benefit shows up under real concurrent traffic. A single request, alone, always waits the full max_wait for nothing — that is expected, and correct.

Try it yourself

Change max_wait from 0.01 to 0.002 and rerun. Latency should drop, throughput should drop too — you are giving the batcher less time to collect company.

What to learn next

  • Continuous batching for LLMs — the version of this idea built for text generation, where requests do not all finish at the same time.
  • Model serving — the service this batcher would sit inside.
  • Latency budgets — deciding how much of your time budget batching is allowed to spend.

Researcher — Mathematics and papers.

The trade-off, made precise

Let $c$ be the fixed cost of one call and $s$ the marginal cost per item. A batch of size $b$ costs $c + sb$, so the amortised cost per item is:

$$\frac{c + sb}{b} = \frac{c}{b} + s$$

As $b \to \infty$, per-item cost approaches $s$ — the fixed cost is fully amortised. The added latency for the request that arrives first in a batch is bounded by $\max_wait$; the request that arrives last pays almost no extra wait but still shares the batch's full compute time, $c + sb$.

This is why batch size and wait time are set together, not independently: a large max_batch with a short max_wait behaves like a small batcher during quiet periods, and only behaves like a large one during bursts.

Where this sits in a serving stack

NVIDIA Triton's dynamic batcher, TensorFlow Serving's batching scheduler, and TorchServe all implement this exact wait-or-fill pattern, configurable per model. The tunables are consistently the same two knobs shown above, sometimes named max_batch_size and max_queue_delay_microseconds.

Papers and references

  • Olston et al., TensorFlow-Serving: Flexible, High-Performance ML Serving, 2017 — arxiv.org/abs/1712.06139
  • Crankshaw et al., Clipper: A Low-Latency Online Prediction Serving System, NSDI 2017 — arxiv.org/abs/1612.03079 — one of the earliest systems papers to formalise adaptive batching against a latency SLO.

What to learn next

  • Continuous batching for LLMs — the version of this idea built for text generation, where requests do not all finish at the same time.
  • Model serving — the service this batcher would sit inside.
  • Latency budgets — deciding how much of your time budget batching is allowed to spend.