Throughput
In one sentence Throughput is how much work a system completes per unit time — requests per second, tokens per second — and it decides your cost per query.
Updated
Throughput is the total amount of work a system completes per unit of time: requests per second, images per minute, tokens per second across all users.
It is the tea stall's daily count, not any single customer's wait. A stall that serves 600 cups an hour has high throughput even if each person queues a while; a boutique café making one elaborate pour-over at a time has splendid per-cup attention and hopeless volume. Which number matters depends on whether you are the customer (latency) or the owner paying rent (throughput) — for serving infrastructure, throughput is what sets cost per request, because the GPU's monthly bill is fixed and throughput decides how many requests share it.
The central lever is batching. A GPU is a parallel machine: processing 32 requests together costs little more time than processing one, so grouping raises throughput enormously. The catch is the trade: requests wait for their batch, and everyone in a batch shares the machine — so throughput tuning nudges individual latency upward. Serving is the art of choosing your point on that curve; LLM stacks like vLLM push the frontier outward with continuous batching and paged KV caches, which is their entire reason for existing.
Units to keep straight in LLM land: tokens/second for one user (generation speed you watch) versus total tokens/second across the server (what you provision by). Capacity planning runs on the second: peak demand ÷ per-GPU throughput = GPUs to buy. Raise throughput without new hardware via quantization, better batching, and shorter prompts; when those run out, scale horizontally.