Continuous batching for LLMs
Continuous batching lets a finished request leave the batch and a new one take its seat immediately, instead of the whole batch waiting for its slowest member.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Continuous batching lets a finished request leave the batch, and a new one join, without anyone waiting on the group's slowest member.
The analogy you have already lived
You have eaten at a busy dosa counter with four stools. In the old style, a fresh batch starts only once all four current customers finish and leave together. That holds even if two of them finished five minutes ago and are sitting there, done. In the good version, the moment one stool empties, the next waiting customer sits down immediately. The counter is never idle while someone else waits outside.
Text generation has exactly this problem. Some replies are one sentence. Some are twenty. Batching them the old way wastes enormous time.
Why it exists
Dynamic batching groups requests and runs them together, which works well when every request takes about the same time. Language models break that assumption. A model writes a reply one token at a time. A token is a small chunk of text the model reads or writes in a single step. Nobody knows in advance how many tokens a given reply will need.
Say four requests enter a batch together, and one needs 200 tokens while the other three need 20. The old style of batching holds all four "seats" busy for the full 200 steps. Three of those seats are producing nothing for 180 of those steps, and no new request can start until every seat is free.
Continuous batching — also called iteration-level scheduling — checks in after every single token, not after every full reply. Any finished request leaves immediately. Any waiting request takes the empty seat immediately. The hardware stays full.
How it works
Old way (batch of 4, wait for the slowest):
seat A: ####-------------------------- (done early, seat wasted)
seat B: ############################## (the slow one -- sets the pace)
seat C: ######---------------------- (done early, seat wasted)
seat D: ############---------------- (done early, seat wasted)
new requests wait outside the WHOLE time
Continuous batching (a seat is refilled the instant it frees up):
seat A: ####[new]#####[new]########### -- always working on something
seat B: ##############################
seat C: ######[new]###[new]#[new]#####
seat D: ############[new]#############A real example you have seen
ChatGPT and similar chat apps serve enormous numbers of conversations at once, with wildly different reply lengths. One person gets a one-line answer next to someone else's ten-paragraph explanation. Serving all of them well, at the same time, on the same hardware, needs exactly this scheduling trick.
The honest part
This only matters once you have more requests than free seats — a quiet server with plenty of spare capacity does not need it. It also needs specific engineering underneath (managing memory for each in-progress reply independently) that a plain batching loop does not need. It is a real piece of systems engineering, not a small tweak.
Remember this
- Old-style batching holds a whole group hostage to its slowest member.
- Continuous batching frees and refills one seat at a time, keeping the hardware full.
- It matters specifically for variable-length work — text generation is the clearest example.
What to learn next
- The KV cache — the per-request memory that makes continuous batching hard, and PagedAttention necessary.
- Concurrency in a Python inference server — how a server juggles many in-flight requests at the code level.
- Time to first token vs tokens per second — the two numbers continuous batching is trying to improve at once.
Developer — Code and libraries.
Setup
No installation needed. This is a scheduling simulation in plain Python — it shows why the technique wins, using ticks of time instead of a real GPU. Running this against a real model needs vLLM or a similar server; that part is called out explicitly below.
Simulating both scheduling policies
import random
B = 4 # GPU slots: how many sequences can be worked on in one tick
def static_batches(lengths):
"""Old-style batching: group requests B at a time. Run each group
until its SLOWEST member finishes, only then start the next group."""
total_ticks = 0
finish_tick = {}
for start in range(0, len(lengths), B):
chunk = lengths[start:start + B]
total_ticks += max(chunk)
for i, length in enumerate(chunk):
finish_tick[start + i] = total_ticks
return total_ticks, finish_tick
def continuous_batches(lengths):
"""New-style batching: a request leaves its slot the instant it is
done, and the next waiting request takes the slot immediately."""
waiting = list(enumerate(lengths)) # (request_id, remaining tokens)
active = {} # slot -> remaining tokens
finish_tick = {}
tick = 0
while waiting or active:
while len(active) < B and waiting:
rid, remaining = waiting.pop(0)
active[rid] = remaining
tick += 1
for rid in list(active):
active[rid] -= 1
if active[rid] == 0:
finish_tick[rid] = tick
del active[rid]
return tick, finish_tick
if __name__ == "__main__":
random.seed(3)
lengths = [random.randint(3, 25) for _ in range(20)] # tokens each reply needs
print("reply lengths (tokens):", lengths)
static_total, static_finish = static_batches(lengths)
cont_total, cont_finish = continuous_batches(lengths)
static_avg = sum(static_finish.values()) / len(static_finish)
cont_avg = sum(cont_finish.values()) / len(cont_finish)
print(f"static batching: {static_total:3d} GPU ticks total, "
f"avg completion {static_avg:5.2f} ticks/request")
print(f"continuous batch: {cont_total:3d} GPU ticks total, "
f"avg completion {cont_avg:5.2f} ticks/request")
print(f"fewer wasted ticks: {static_total - cont_total} "
f"({(1 - cont_total/static_total)*100:.0f}% less GPU time for the same work)")reply lengths (tokens): [10, 21, 20, 7, 14, 22, 18, 23, 21, 5, 22, 3, 18, 11, 20, 10, 9, 25, 18, 20] static batching: 111 GPU ticks total, avg completion 65.60 ticks/request continuous batch: 84 GPU ticks total, avg completion 46.40 ticks/request fewer wasted ticks: 27 (24% less GPU time for the same work)
This output is deterministic — the random seed is fixed, so rerunning gives the exact same numbers. That is a property of the simulation, not a claim about real hardware. A real GPU's numbers depend on model size, sequence lengths, and memory bandwidth, and need to be measured on real hardware to trust.
Line-by-line walkthrough
static_batches groups requests into fixed chunks of size B. Every chunk is charged the cost of its slowest member — max(chunk) — because that is genuinely how old-style batching behaves: nothing new starts until every seat is free.
continuous_batches runs a tick-by-tick loop. Every tick, it tops up active from waiting before doing any work, then advances every active request by one tick. A request that hits zero remaining tokens is removed and reported as finished, on that exact tick — freeing its slot for the very next tick.
Common mistakes
Thinking continuous batching is only about speed. It also drastically shortens the wait before your reply even starts, because new requests are no longer stuck outside a full batch — see the avg completion numbers above.
Assuming it needs no extra memory management. Each in-progress reply needs its own KV cache — memory holding what the model has already computed for that conversation. Continuous batching means that memory keeps growing and shrinking per seat. That churn is exactly the problem PagedAttention was built to manage.
Comparing a hand-rolled server against a mature one. vLLM and NVIDIA's TGI implement continuous batching with years of tuning underneath. Reimplementing the scheduling loop above is for understanding it, not for shipping it.
Try it yourself
Change B from 4 to 8 and rerun. Both totals should drop, since more seats mean fewer wasted ticks. But the relative saving from continuous batching should shrink too — a bigger static batch already hides more of its own waste.
What to learn next
- The KV cache — the per-request memory that makes continuous batching hard, and PagedAttention necessary.
- Concurrency in a Python inference server — how a server juggles many in-flight requests at the code level.
- Time to first token vs tokens per second — the two numbers continuous batching is trying to improve at once.
Researcher — Mathematics and papers.
Iteration-level scheduling
Orca (Yu et al., 2022) introduced iteration-level scheduling: rather than scheduling at the request level (admit a batch, run it to completion, repeat), the scheduler makes an admission decision at every iteration of the model's forward pass. A finished sequence is evicted and a queued one admitted within the same iteration boundary, matching exactly the simulation above.
PagedAttention
Continuous batching creates a new problem: each active sequence's KV cache grows every step, by an unpredictable final amount, and naive contiguous allocation fragments GPU memory badly under this churn. vLLM (Kwon et al., 2023) introduced PagedAttention, which manages the KV cache in fixed-size, non-contiguous blocks — analogous to virtual memory paging in an operating system. This removed the fragmentation that had capped achievable batch sizes, and is now close to a de facto standard: Hugging Face TGI and most current serving stacks use the same idea.
Throughput reported in the literature
Reported gains over prior static-batching systems are commonly cited in the 2–4× range for realistic chat-style workloads with mixed reply lengths (Kwon et al., 2023). This figure comes from the paper's own benchmarks on specific hardware and workloads — treat it as a plausible ballpark, not a guarantee for any particular deployment, and measure your own workload before relying on it.
Papers
- Yu et al., Orca: A Distributed Serving System for Transformer-Based Generative Models, OSDI 2022
- Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM), SOSP 2023 — arxiv.org/abs/2309.06180
- Yu et al.'s Orca and the vLLM paper together form the standard citation pair for this technique; most later serving-systems papers compare against one or both.
What to learn next
- The KV cache — the per-request memory that makes continuous batching hard, and PagedAttention necessary.
- Concurrency in a Python inference server — how a server juggles many in-flight requests at the code level.
- Time to first token vs tokens per second — the two numbers continuous batching is trying to improve at once.