Batching
In one sentence Batching groups multiple inputs and processes them together in one pass, the key to keeping GPUs busy in both training and serving.
Updated
Batching means processing many inputs together in one pass through the model, instead of one at a time.
A school bus and a car both take children to school. The car makes thirty trips; the bus makes one. The road time is nearly identical either way — the vehicle is going regardless — so seats filled per trip is almost free efficiency. A GPU is the bus: it is a massively parallel machine, and the fixed costs of a pass (loading weights from memory, launching compute kernels) get paid once whether one input rides along or sixty-four. Much of a model's arithmetic is memory-bound, so those weights fetched once serve every passenger.
The term spans two worlds. In training, batch-size is a hyperparameter affecting both speed and learning dynamics. In serving, batching is a systems trick: collect the requests arriving now, run them together, split the results — multiplying throughput at some cost to individual latency, since requests briefly wait for companions.
LLM serving complicated this beautifully. Requests generate different numbers of tokens, so naive ("static") batches finish at the speed of their longest member while finished slots sit idle. The fix, continuous batching, operates at the token level: every generation step, completed sequences exit the batch and queued requests join mid-flight. Seats are refilled at every bus stop, not once per route. This single idea is the core of modern LLM servers such as vLLM and TGI, and it is why they achieve several-fold higher throughput than naive serving at similar latency.
Where to go next
- Full lesson: vLLM
- Related terms: throughput, latency, batch-size, gpu