vLLM
vLLM is a server that runs one language model on a GPU for many users at once, by never letting the graphics card sit idle waiting.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
vLLM is a server that lets one copy of a language model answer many people at the same time, quickly.
The analogy you have already lived
Picture two food stalls at a station.
The first cook takes one order, finishes it completely, hands it over, and only then looks at you. The queue behind him grows and his second burner sits cold the whole time.
The second cook has four pans going. The moment one dish leaves a pan, the next order goes straight in. Same cook, same stove, four times the plates served.
Ollama is the first cook, and that is the right design for one person on a laptop. vLLM is the second cook.
Why it exists
A graphics card is expensive and rented by the hour. Every second it spends idle is money burned.
Here is the awkward truth about answering with a language model. The model writes one word, then reads everything again to write the next word. Between words, most of the card does nothing useful.
The old fix was static batching: collect eight questions, run all eight together, send back eight answers. It helps, and it has an ugly flaw. The batch is only finished when the slowest answer finishes. Seven users who asked short questions sit and wait for the one who asked for an essay.
vLLM's answer is continuous batching. A finished answer leaves the batch immediately, and a waiting request takes its place in that same instant. Nobody waits for a stranger's essay.
There was a second waste, and it was bigger. While the model writes, it keeps notes about the conversation so far. Older servers reserved room for the longest conversation the model could ever have, for every single user, up front. Most of that room was never used.
vLLM hands out that memory in small pieces, only as the conversation actually grows. Think of a hostel giving each student one locker at a time. Nobody is assigned a whole cupboard on day one.
The technique is called paged attention. It is the reason vLLM fits several times more users on the same card.
How it works
user A ──┐
user B ──┤ +--------------------------+
user C ──┼────> | vLLM server |
user D ──┘ | |
| ONE copy of the model |
| on ONE graphics card |
| |
| A finishes -> A leaves |
| E is waiting -> E joins |
+--------------------------+
|
answers stream back,
word by wordOne model. One card. A queue that never stalls.
Where you have already met it
Every AI chat product you have used serves thousands of people from a shared pool of graphics cards. vLLM, or something built on the same two ideas, is doing that work.
You notice it as the reason answers start appearing in under a second even when millions of people are online.
The honest part
You need an NVIDIA graphics card. Not a preference — a requirement for the normal install. A laptop without one cannot run this.
You need Linux. vLLM does not install on Windows directly. Windows users go through WSL2, which is Linux running inside Windows.
It is the wrong tool for one person. For your own laptop, for learning, for a demo, Ollama is easier. It installs in a minute. Reach for vLLM when many people will use your model at once.
Renting a card costs real money. A modest cloud GPU runs from roughly 25 to 60 rupees an hour. A full month of that is a serious bill for a student. Google Colab's free tier gives you a card for a few hours at a time. That is enough to try everything on this page.
Remember this
- vLLM serves many users from one model on one card, without idle time.
- Continuous batching means a finished answer leaves the queue at once.
- Paged attention hands out memory in small pieces, so far more users fit.
What to learn next
- Model deployment — where to put this server without going broke.
- Ollama — the single-user version, for your own machine.
- Context windows — why that number drives every memory decision here.
Developer — Code and libraries.
Before you install: does your machine qualify?
vLLM's standard build needs Linux and an NVIDIA GPU with CUDA. There are builds for AMD ROCm, Intel and CPU, but they are compiled from source and are not the beginner path.
If you have no GPU, the honest options are:
- Google Colab — free tier gives a T4 card. Enough for a 1.5B model. Sessions end without warning.
- Kaggle Notebooks — free weekly GPU hours, longer sessions than Colab.
- A rented cloud GPU — pay by the hour, shut it down when you sleep.
Everything below runs on a free Colab T4.
Install
pip install vllmThat pulls PyTorch with CUDA and is a multi-gigabyte download. Do it on Wi-Fi, not mobile data.
Serve a model
vllm serve Qwen/Qwen2.5-1.5B-Instruct \
--max-model-len 4096 \
--gpu-memory-utilization 0.90 \
--dtype halfThree flags, and each one prevents a specific failure.
--max-model-len 4096 caps the conversation length. The default is whatever the model declares, often 32,768 or more, and vLLM reserves memory for it at startup. This flag is the difference between starting and crashing.
--gpu-memory-utilization 0.90 tells vLLM it may use 90% of the card. Lower it if something else is sharing the GPU.
--dtype half forces float16. On older cards — the T4 in free Colab included — bfloat16 is unsupported, and without this flag you get:
ValueError: Bfloat16 is only supported on GPUs with compute capability of at least 8.0. Your Tesla T4 GPU has compute capability 7.5. You can use float16 instead by explicitly setting the `dtype` flag in CLI, for example: --dtype=half.
The message names your card, so the wording differs. On an A100, L4, or any RTX 30-series or newer, leave --dtype off and let vLLM choose.
Talk to it
The server speaks the OpenAI API format, so every OpenAI client library works against it unchanged.
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen2.5-1.5B-Instruct",
"messages": [{"role": "user", "content": "Name three uses of a vector database."}],
"max_tokens": 120
}'Or from Python:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")
response = client.chat.completions.create(
model="Qwen/Qwen2.5-1.5B-Instruct",
messages=[{"role": "user", "content": "Name three uses of a vector database."}],
max_tokens=120,
)
print(response.choices[0].message.content)api_key is required by the client library and ignored by a local server. Put anything there.
No output block for these two, deliberately. The wording changes with every model, every version and every run. An invented sample would teach you to expect something that will not happen.
Working out whether a model fits, before you rent anything
This is the calculation that decides your whole hardware budget, and it is arithmetic you can run on any laptop right now. No GPU, no vLLM, no downloads.
GIB = 1024 ** 3
GPU_GIB = 16.0 # one NVIDIA T4 / RTX 4060 Ti class card
UTILISATION = 0.90 # vLLM's default --gpu-memory-utilization
def kv_bytes_per_token(layers, kv_heads, head_dim, bytes_per_value=2):
# The 2 is because a key tensor AND a value tensor are cached, per layer.
return 2 * layers * kv_heads * head_dim * bytes_per_value
MODELS = {
# name: (layers, kv_heads, head_dim, fp16 weight GiB)
"Qwen2.5-1.5B-Instruct": (28, 2, 128, 2.9),
"Llama-3.1-8B-Instruct": (32, 8, 128, 15.0),
}
for name, (layers, kv_heads, head_dim, weights) in MODELS.items():
per_token = kv_bytes_per_token(layers, kv_heads, head_dim)
kv_budget = GPU_GIB * UTILISATION - weights
print(name)
print(f" KV cache per token : {per_token / 1024:.0f} KiB")
print(f" weights in fp16 : {weights:.1f} GiB")
if kv_budget <= 0:
print(f" KV budget : {kv_budget:.1f} GiB -> does not fit, quantise it")
else:
tokens = int(kv_budget * GIB / per_token)
print(f" KV budget : {kv_budget:.1f} GiB -> {tokens:,} cached tokens")
print(f" roughly {tokens // 4096} users at once, 4096 tokens each")
print()Qwen2.5-1.5B-Instruct KV cache per token : 28 KiB weights in fp16 : 2.9 GiB KV budget : 11.5 GiB -> 430,665 cached tokens roughly 105 users at once, 4096 tokens each Llama-3.1-8B-Instruct KV cache per token : 128 KiB weights in fp16 : 15.0 GiB KV budget : -0.6 GiB -> does not fit, quantise it
What that output is telling you
Llama-3.1-8B does not fit on a 16 GB card in full precision. The weights alone are 15 GiB, leaving nothing for the conversation cache. vLLM will start, then fail while allocating. The fix is a quantized checkpoint — an AWQ or FP8 build — which roughly halves the weights.
The 1.5B model has room for about a hundred concurrent users. That number is the honest capacity of the machine, and it came from four numbers you can read off a model's config file on Hugging Face: num_hidden_layers, num_key_value_heads, hidden_size / num_attention_heads, and the parameter count.
Notice the gap between 28 KiB and 128 KiB per token. Both models use 128-dimension heads. The 8B model caches four times as many key/value heads per layer, and has more layers. Cache cost is not proportional to model size, and you cannot guess it.
Offline batch mode, when there is no server
If you have a pile of work and no users waiting, skip the server entirely.
from vllm import LLM, SamplingParams
llm = LLM(model="Qwen/Qwen2.5-1.5B-Instruct", max_model_len=2048, dtype="half")
params = SamplingParams(temperature=0.0, max_tokens=64)
prompts = [
"Summarise in one line: the train was late by two hours.",
"Summarise in one line: the parcel arrived damaged.",
]
for output in llm.generate(prompts, params):
print(output.outputs[0].text.strip())This is the fastest way to process ten thousand rows. vLLM batches them internally, and the graphics card stays busy from start to finish. Generated text varies by model and version, so no output block here either.
Common mistakes
Out of memory at startup, not during use. vLLM claims its memory up front. An OOM in the first ten seconds means --max-model-len is too high, or the model is too big for the card. It is almost never a leak.
Setting --gpu-memory-utilization 1.0. The card also holds the CUDA context and fragmentation slack. Above roughly 0.95 you get crashes that look random. Leave it at 0.90.
Benchmarking with one request at a time. vLLM's advantage is throughput under load. Send one request and it can look slower than Ollama. Use the bundled vllm bench serve tool, or send fifty parallel requests yourself, before drawing any conclusion.
Forgetting the Hugging Face token for gated models. Llama checkpoints need an accepted licence. Without HF_TOKEN set, the download fails with HTTP 401.
Running it on your laptop for a class demo. If one person is using the model, Ollama starts faster, installs in a minute and works on Windows. Use the right tool.
Try it yourself
Run fit.py after changing GPU_GIB to 24 — an RTX 3090 or 4090. Then look up config.json for a model you actually want to serve, read the four numbers out of it, and add it to the MODELS dictionary. That habit, checked before you rent anything, will save you more money than any other five minutes in this section.
What to learn next
- Model deployment — where to put this server without going broke.
- Ollama — the single-user version, for your own machine.
- Context windows — why that number drives every memory decision here.
Researcher — Mathematics and papers.
PagedAttention
Kwon et al. (2023) identified that pre-vLLM serving systems allocated the KV cache as one contiguous tensor per sequence, sized to the maximum generation length. They measured 60–80% of that memory as waste, in three forms: internal fragmentation (reserved but never generated), external fragmentation (unusable gaps between allocations), and reservation (held for future tokens of a live sequence).
PagedAttention borrows virtual memory paging. The KV cache is partitioned into fixed-size blocks (typically 16 tokens). A per-sequence block table maps logical token positions to arbitrary physical blocks, so a sequence's cache need not be contiguous. The attention kernel is rewritten to gather across blocks.
Consequences:
- Waste is bounded by one partially-filled block per sequence, so under 4% instead of 60–80%.
- Copy-on-write block sharing becomes possible. Parallel samples from one prompt, or beam search branches, share prompt blocks with a reference count, and fork only on divergence.
- The reported end-to-end result is 2–4× throughput at equal latency versus HuggingFace TGI and FasterTransformer of the time.
Continuous batching
Also called iteration-level scheduling; introduced as Orca (Yu et al., 2022). The scheduler runs per decoding step, not per request. After each forward pass, finished sequences are evicted and queued sequences admitted.
Static batching wastes capacity proportional to the variance in output length. Since output lengths are long-tailed in practice, that variance is large, and this is where most of the gain comes from.
Preemption matters at saturation. When the block pool is exhausted, vLLM either swaps blocks to host memory or recomputes them on resumption. Recomputation is usually cheaper, because prefill is compute-bound and PCIe transfer is not.
Prefill and decode are different workloads
| Phase | Parallelism | Bound by | Arithmetic intensity |
|---|---|---|---|
| Prefill | all prompt tokens at once | compute (FLOPs) | high |
| Decode | one token per sequence | memory bandwidth | ~1 op per weight byte |
Mixing them in one batch means a long prefill blocks every decode step in flight, producing inter-token latency spikes. Chunked prefill (Agrawal et al., Sarathi, 2023) splits a long prefill into token-budgeted chunks and co-schedules each chunk with ongoing decodes, trading a small prefill slowdown for a large tail-latency improvement. It is on by default in current vLLM.
The alternative is disaggregated serving (Zhong et al., DistServe, 2024): run prefill and decode on separate GPU pools and ship the KV cache between them. This wins when the two phases have genuinely different SLOs, at the cost of an interconnect dependency.
Prefix caching
Requests that share a prefix — a long system prompt, a few-shot block, a shared document — can share KV blocks. vLLM hashes block contents and reuses matching blocks across requests (--enable-prefix-caching, default on in V1).
For a workload with a 2,000-token shared system prompt and 100-token user turns, this removes roughly 95% of prefill compute. It is frequently the largest single win available, and it costs one flag.
Capacity, via Little's law
L = λ · WL is mean concurrency, λ the arrival rate, W the mean time in system. Concurrency is capped by the block pool:
L_max = total_kv_bytes / (kv_bytes_per_token × mean_sequence_tokens)Substituting gives sustainable throughput. The practical reading: halving kv_bytes_per_token — through GQA, FP8 KV cache, or a smaller --max-model-len — doubles concurrency at fixed latency. This is why grouped-query attention (Ainslie et al., 2023) changed serving economics more than any kernel optimisation.
Note the failure mode when λ exceeds capacity: queueing delay grows without bound, and p99 latency degrades far faster than mean latency. Provision on p99, and measure with a load generator, not with time curl.
Parallelism
- Tensor parallelism (
--tensor-parallel-size) shards each weight matrix across GPUs. Requires an all-reduce per layer, so it wants NVLink; over PCIe the collective cost dominates quickly. - Pipeline parallelism (
--pipeline-parallel-size) shards by layer. Cheaper interconnect, but introduces bubbles unless microbatched. - Expert parallelism for mixture-of-experts models, routing tokens to expert shards.
Rule of thumb: use the smallest tensor-parallel degree that fits the weights plus a workable KV budget. Sharding beyond necessity costs latency for nothing.
Quantization in serving
- FP8 (E4M3) on Hopper and later — near-lossless, with hardware-accelerated matmul. The default choice where available.
- AWQ (Lin et al., 2023) and GPTQ (Frantar et al., 2022) — 4-bit weight-only. Weights shrink 4×; activations stay 16-bit. Excellent for memory-bound decode, less useful for compute-bound prefill.
- FP8 KV cache (
--kv-cache-dtype fp8) — halves the cache, roughly doubling concurrency. Usually a better return than quantizing weights further.
Papers
- Kwon et al., Efficient Memory Management for LLM Serving with PagedAttention, SOSP 2023 — arxiv.org/abs/2309.06180
- Yu et al., Orca: A Distributed Serving System for Transformer-Based Generative Models, OSDI 2022
- Agrawal et al., SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills, 2023 — arxiv.org/abs/2308.16369
- Zhong et al., DistServe, OSDI 2024 — arxiv.org/abs/2401.09670
- Dao et al., FlashAttention-2, 2023 — arxiv.org/abs/2307.08691
- Pope et al., Efficiently Scaling Transformer Inference, 2022 — arxiv.org/abs/2211.05102
What to learn next
- Model deployment — where to put this server without going broke.
- Ollama — the single-user version, for your own machine.
- Context windows — why that number drives every memory decision here.