Batching and Concurrency

Continuous batching for LLMs

Continuous batching lets a finished request leave the batch and a new one take its seat immediately, instead of the whole batch waiting for its slowest member.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Continuous batching lets a finished request leave the batch, and a new one join, without anyone waiting on the group's slowest member.

The analogy you have already lived

You have eaten at a busy dosa counter with four stools. In the old style, a fresh batch starts only once all four current customers finish and leave together. That holds even if two of them finished five minutes ago and are sitting there, done. In the good version, the moment one stool empties, the next waiting customer sits down immediately. The counter is never idle while someone else waits outside.

Text generation has exactly this problem. Some replies are one sentence. Some are twenty. Batching them the old way wastes enormous time.

Why it exists

Dynamic batching groups requests and runs them together, which works well when every request takes about the same time. Language models break that assumption. A model writes a reply one token at a time. A token is a small chunk of text the model reads or writes in a single step. Nobody knows in advance how many tokens a given reply will need.

Say four requests enter a batch together, and one needs 200 tokens while the other three need 20. The old style of batching holds all four "seats" busy for the full 200 steps. Three of those seats are producing nothing for 180 of those steps, and no new request can start until every seat is free.

Continuous batching — also called iteration-level scheduling — checks in after every single token, not after every full reply. Any finished request leaves immediately. Any waiting request takes the empty seat immediately. The hardware stays full.

How it works

Old way (batch of 4, wait for the slowest):
  seat A: ####--------------------------  (done early, seat wasted)
  seat B: ##############################  (the slow one -- sets the pace)
  seat C: ######----------------------    (done early, seat wasted)
  seat D: ############----------------    (done early, seat wasted)
                                            new requests wait outside the WHOLE time

Continuous batching (a seat is refilled the instant it frees up):
  seat A: ####[new]#####[new]###########   -- always working on something
  seat B: ##############################
  seat C: ######[new]###[new]#[new]#####
  seat D: ############[new]#############

A real example you have seen

ChatGPT and similar chat apps serve enormous numbers of conversations at once, with wildly different reply lengths. One person gets a one-line answer next to someone else's ten-paragraph explanation. Serving all of them well, at the same time, on the same hardware, needs exactly this scheduling trick.

The honest part

This only matters once you have more requests than free seats — a quiet server with plenty of spare capacity does not need it. It also needs specific engineering underneath (managing memory for each in-progress reply independently) that a plain batching loop does not need. It is a real piece of systems engineering, not a small tweak.

Remember this

  • Old-style batching holds a whole group hostage to its slowest member.
  • Continuous batching frees and refills one seat at a time, keeping the hardware full.
  • It matters specifically for variable-length work — text generation is the clearest example.

What to learn next

Developer — Code and libraries.

Setup

No installation needed. This is a scheduling simulation in plain Python — it shows why the technique wins, using ticks of time instead of a real GPU. Running this against a real model needs vLLM or a similar server; that part is called out explicitly below.

Simulating both scheduling policies

continuous_batching.py
import random

B = 4  # GPU slots: how many sequences can be worked on in one tick


def static_batches(lengths):
    """Old-style batching: group requests B at a time. Run each group
    until its SLOWEST member finishes, only then start the next group."""
    total_ticks = 0
    finish_tick = {}
    for start in range(0, len(lengths), B):
        chunk = lengths[start:start + B]
        total_ticks += max(chunk)
        for i, length in enumerate(chunk):
            finish_tick[start + i] = total_ticks
    return total_ticks, finish_tick


def continuous_batches(lengths):
    """New-style batching: a request leaves its slot the instant it is
    done, and the next waiting request takes the slot immediately."""
    waiting = list(enumerate(lengths))   # (request_id, remaining tokens)
    active = {}                           # slot -> remaining tokens
    finish_tick = {}
    tick = 0
    while waiting or active:
        while len(active) < B and waiting:
            rid, remaining = waiting.pop(0)
            active[rid] = remaining
        tick += 1
        for rid in list(active):
            active[rid] -= 1
            if active[rid] == 0:
                finish_tick[rid] = tick
                del active[rid]
    return tick, finish_tick


if __name__ == "__main__":
    random.seed(3)
    lengths = [random.randint(3, 25) for _ in range(20)]  # tokens each reply needs
    print("reply lengths (tokens):", lengths)

    static_total, static_finish = static_batches(lengths)
    cont_total, cont_finish = continuous_batches(lengths)

    static_avg = sum(static_finish.values()) / len(static_finish)
    cont_avg = sum(cont_finish.values()) / len(cont_finish)

    print(f"static  batching: {static_total:3d} GPU ticks total, "
          f"avg completion {static_avg:5.2f} ticks/request")
    print(f"continuous batch: {cont_total:3d} GPU ticks total, "
          f"avg completion {cont_avg:5.2f} ticks/request")
    print(f"fewer wasted ticks: {static_total - cont_total} "
          f"({(1 - cont_total/static_total)*100:.0f}% less GPU time for the same work)")
Output
reply lengths (tokens): [10, 21, 20, 7, 14, 22, 18, 23, 21, 5, 22, 3, 18, 11, 20, 10, 9, 25, 18, 20]
static  batching: 111 GPU ticks total, avg completion 65.60 ticks/request
continuous batch:  84 GPU ticks total, avg completion 46.40 ticks/request
fewer wasted ticks: 27 (24% less GPU time for the same work)

This output is deterministic — the random seed is fixed, so rerunning gives the exact same numbers. That is a property of the simulation, not a claim about real hardware. A real GPU's numbers depend on model size, sequence lengths, and memory bandwidth, and need to be measured on real hardware to trust.

Line-by-line walkthrough

static_batches groups requests into fixed chunks of size B. Every chunk is charged the cost of its slowest member — max(chunk) — because that is genuinely how old-style batching behaves: nothing new starts until every seat is free.

continuous_batches runs a tick-by-tick loop. Every tick, it tops up active from waiting before doing any work, then advances every active request by one tick. A request that hits zero remaining tokens is removed and reported as finished, on that exact tick — freeing its slot for the very next tick.

Common mistakes

Thinking continuous batching is only about speed. It also drastically shortens the wait before your reply even starts, because new requests are no longer stuck outside a full batch — see the avg completion numbers above.

Assuming it needs no extra memory management. Each in-progress reply needs its own KV cache — memory holding what the model has already computed for that conversation. Continuous batching means that memory keeps growing and shrinking per seat. That churn is exactly the problem PagedAttention was built to manage.

Comparing a hand-rolled server against a mature one. vLLM and NVIDIA's TGI implement continuous batching with years of tuning underneath. Reimplementing the scheduling loop above is for understanding it, not for shipping it.

Try it yourself

Change B from 4 to 8 and rerun. Both totals should drop, since more seats mean fewer wasted ticks. But the relative saving from continuous batching should shrink too — a bigger static batch already hides more of its own waste.

What to learn next

Researcher — Mathematics and papers.

Iteration-level scheduling

Orca (Yu et al., 2022) introduced iteration-level scheduling: rather than scheduling at the request level (admit a batch, run it to completion, repeat), the scheduler makes an admission decision at every iteration of the model's forward pass. A finished sequence is evicted and a queued one admitted within the same iteration boundary, matching exactly the simulation above.

PagedAttention

Continuous batching creates a new problem: each active sequence's KV cache grows every step, by an unpredictable final amount, and naive contiguous allocation fragments GPU memory badly under this churn. vLLM (Kwon et al., 2023) introduced PagedAttention, which manages the KV cache in fixed-size, non-contiguous blocks — analogous to virtual memory paging in an operating system. This removed the fragmentation that had capped achievable batch sizes, and is now close to a de facto standard: Hugging Face TGI and most current serving stacks use the same idea.

Throughput reported in the literature

Reported gains over prior static-batching systems are commonly cited in the 2–4× range for realistic chat-style workloads with mixed reply lengths (Kwon et al., 2023). This figure comes from the paper's own benchmarks on specific hardware and workloads — treat it as a plausible ballpark, not a guarantee for any particular deployment, and measure your own workload before relying on it.

Papers

  • Yu et al., Orca: A Distributed Serving System for Transformer-Based Generative Models, OSDI 2022
  • Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM), SOSP 2023 — arxiv.org/abs/2309.06180
  • Yu et al.'s Orca and the vLLM paper together form the standard citation pair for this technique; most later serving-systems papers compare against one or both.

What to learn next