LLM Development

vLLM

vLLM is a server that runs one language model on a GPU for many users at once, by never letting the graphics card sit idle waiting.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. Where you have already met it
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

vLLM is a server that lets one copy of a language model answer many people at the same time, quickly.

The analogy you have already lived

Picture two food stalls at a station.

The first cook takes one order, finishes it completely, hands it over, and only then looks at you. The queue behind him grows and his second burner sits cold the whole time.

The second cook has four pans going. The moment one dish leaves a pan, the next order goes straight in. Same cook, same stove, four times the plates served.

Ollama is the first cook, and that is the right design for one person on a laptop. vLLM is the second cook.

Why it exists

A graphics card is expensive and rented by the hour. Every second it spends idle is money burned.

Here is the awkward truth about answering with a language model. The model writes one word, then reads everything again to write the next word. Between words, most of the card does nothing useful.

The old fix was static batching: collect eight questions, run all eight together, send back eight answers. It helps, and it has an ugly flaw. The batch is only finished when the slowest answer finishes. Seven users who asked short questions sit and wait for the one who asked for an essay.

vLLM's answer is continuous batching. A finished answer leaves the batch immediately, and a waiting request takes its place in that same instant. Nobody waits for a stranger's essay.

There was a second waste, and it was bigger. While the model writes, it keeps notes about the conversation so far. Older servers reserved room for the longest conversation the model could ever have, for every single user, up front. Most of that room was never used.

vLLM hands out that memory in small pieces, only as the conversation actually grows. Think of a hostel giving each student one locker at a time. Nobody is assigned a whole cupboard on day one.

The technique is called paged attention. It is the reason vLLM fits several times more users on the same card.

How it works

   user A ──┐
   user B ──┤      +--------------------------+
   user C ──┼────> |  vLLM server             |
   user D ──┘      |                          |
                   |  ONE copy of the model   |
                   |  on ONE graphics card    |
                   |                          |
                   |  A finishes -> A leaves  |
                   |  E is waiting -> E joins |
                   +--------------------------+
                              |
                     answers stream back,
                     word by word

One model. One card. A queue that never stalls.

Where you have already met it

Every AI chat product you have used serves thousands of people from a shared pool of graphics cards. vLLM, or something built on the same two ideas, is doing that work.

You notice it as the reason answers start appearing in under a second even when millions of people are online.

The honest part

You need an NVIDIA graphics card. Not a preference — a requirement for the normal install. A laptop without one cannot run this.

You need Linux. vLLM does not install on Windows directly. Windows users go through WSL2, which is Linux running inside Windows.

It is the wrong tool for one person. For your own laptop, for learning, for a demo, Ollama is easier. It installs in a minute. Reach for vLLM when many people will use your model at once.

Renting a card costs real money. A modest cloud GPU runs from roughly 25 to 60 rupees an hour. A full month of that is a serious bill for a student. Google Colab's free tier gives you a card for a few hours at a time. That is enough to try everything on this page.

Remember this

  • vLLM serves many users from one model on one card, without idle time.
  • Continuous batching means a finished answer leaves the queue at once.
  • Paged attention hands out memory in small pieces, so far more users fit.

What to learn next

  • Model deployment — where to put this server without going broke.
  • Ollama — the single-user version, for your own machine.
  • Context windows — why that number drives every memory decision here.

Developer — Code and libraries.

Before you install: does your machine qualify?

vLLM's standard build needs Linux and an NVIDIA GPU with CUDA. There are builds for AMD ROCm, Intel and CPU, but they are compiled from source and are not the beginner path.

If you have no GPU, the honest options are:

  • Google Colab — free tier gives a T4 card. Enough for a 1.5B model. Sessions end without warning.
  • Kaggle Notebooks — free weekly GPU hours, longer sessions than Colab.
  • A rented cloud GPU — pay by the hour, shut it down when you sleep.

Everything below runs on a free Colab T4.

Install

bash
pip install vllm

That pulls PyTorch with CUDA and is a multi-gigabyte download. Do it on Wi-Fi, not mobile data.

Serve a model

bash
vllm serve Qwen/Qwen2.5-1.5B-Instruct \
  --max-model-len 4096 \
  --gpu-memory-utilization 0.90 \
  --dtype half

Three flags, and each one prevents a specific failure.

--max-model-len 4096 caps the conversation length. The default is whatever the model declares, often 32,768 or more, and vLLM reserves memory for it at startup. This flag is the difference between starting and crashing.

--gpu-memory-utilization 0.90 tells vLLM it may use 90% of the card. Lower it if something else is sharing the GPU.

--dtype half forces float16. On older cards — the T4 in free Colab included — bfloat16 is unsupported, and without this flag you get:

Output
ValueError: Bfloat16 is only supported on GPUs with compute capability of at least 8.0. Your Tesla T4 GPU has compute capability 7.5. You can use float16 instead by explicitly setting the `dtype` flag in CLI, for example: --dtype=half.

The message names your card, so the wording differs. On an A100, L4, or any RTX 30-series or newer, leave --dtype off and let vLLM choose.

Talk to it

The server speaks the OpenAI API format, so every OpenAI client library works against it unchanged.

bash
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen2.5-1.5B-Instruct",
    "messages": [{"role": "user", "content": "Name three uses of a vector database."}],
    "max_tokens": 120
  }'

Or from Python:

client.py
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")

response = client.chat.completions.create(
    model="Qwen/Qwen2.5-1.5B-Instruct",
    messages=[{"role": "user", "content": "Name three uses of a vector database."}],
    max_tokens=120,
)
print(response.choices[0].message.content)

api_key is required by the client library and ignored by a local server. Put anything there.

No output block for these two, deliberately. The wording changes with every model, every version and every run. An invented sample would teach you to expect something that will not happen.

Working out whether a model fits, before you rent anything

This is the calculation that decides your whole hardware budget, and it is arithmetic you can run on any laptop right now. No GPU, no vLLM, no downloads.

fit.py
GIB = 1024 ** 3
GPU_GIB = 16.0          # one NVIDIA T4 / RTX 4060 Ti class card
UTILISATION = 0.90      # vLLM's default --gpu-memory-utilization


def kv_bytes_per_token(layers, kv_heads, head_dim, bytes_per_value=2):
    # The 2 is because a key tensor AND a value tensor are cached, per layer.
    return 2 * layers * kv_heads * head_dim * bytes_per_value


MODELS = {
    # name:                   (layers, kv_heads, head_dim, fp16 weight GiB)
    "Qwen2.5-1.5B-Instruct":  (28, 2, 128, 2.9),
    "Llama-3.1-8B-Instruct":  (32, 8, 128, 15.0),
}

for name, (layers, kv_heads, head_dim, weights) in MODELS.items():
    per_token = kv_bytes_per_token(layers, kv_heads, head_dim)
    kv_budget = GPU_GIB * UTILISATION - weights
    print(name)
    print(f"  KV cache per token : {per_token / 1024:.0f} KiB")
    print(f"  weights in fp16    : {weights:.1f} GiB")
    if kv_budget <= 0:
        print(f"  KV budget          : {kv_budget:.1f} GiB -> does not fit, quantise it")
    else:
        tokens = int(kv_budget * GIB / per_token)
        print(f"  KV budget          : {kv_budget:.1f} GiB -> {tokens:,} cached tokens")
        print(f"  roughly {tokens // 4096} users at once, 4096 tokens each")
    print()
Output
Qwen2.5-1.5B-Instruct
  KV cache per token : 28 KiB
  weights in fp16    : 2.9 GiB
  KV budget          : 11.5 GiB -> 430,665 cached tokens
  roughly 105 users at once, 4096 tokens each

Llama-3.1-8B-Instruct
  KV cache per token : 128 KiB
  weights in fp16    : 15.0 GiB
  KV budget          : -0.6 GiB -> does not fit, quantise it

What that output is telling you

Llama-3.1-8B does not fit on a 16 GB card in full precision. The weights alone are 15 GiB, leaving nothing for the conversation cache. vLLM will start, then fail while allocating. The fix is a quantized checkpoint — an AWQ or FP8 build — which roughly halves the weights.

The 1.5B model has room for about a hundred concurrent users. That number is the honest capacity of the machine, and it came from four numbers you can read off a model's config file on Hugging Face: num_hidden_layers, num_key_value_heads, hidden_size / num_attention_heads, and the parameter count.

Notice the gap between 28 KiB and 128 KiB per token. Both models use 128-dimension heads. The 8B model caches four times as many key/value heads per layer, and has more layers. Cache cost is not proportional to model size, and you cannot guess it.

Offline batch mode, when there is no server

If you have a pile of work and no users waiting, skip the server entirely.

batch.py
from vllm import LLM, SamplingParams

llm = LLM(model="Qwen/Qwen2.5-1.5B-Instruct", max_model_len=2048, dtype="half")
params = SamplingParams(temperature=0.0, max_tokens=64)

prompts = [
    "Summarise in one line: the train was late by two hours.",
    "Summarise in one line: the parcel arrived damaged.",
]

for output in llm.generate(prompts, params):
    print(output.outputs[0].text.strip())

This is the fastest way to process ten thousand rows. vLLM batches them internally, and the graphics card stays busy from start to finish. Generated text varies by model and version, so no output block here either.

Common mistakes

Out of memory at startup, not during use. vLLM claims its memory up front. An OOM in the first ten seconds means --max-model-len is too high, or the model is too big for the card. It is almost never a leak.

Setting --gpu-memory-utilization 1.0. The card also holds the CUDA context and fragmentation slack. Above roughly 0.95 you get crashes that look random. Leave it at 0.90.

Benchmarking with one request at a time. vLLM's advantage is throughput under load. Send one request and it can look slower than Ollama. Use the bundled vllm bench serve tool, or send fifty parallel requests yourself, before drawing any conclusion.

Forgetting the Hugging Face token for gated models. Llama checkpoints need an accepted licence. Without HF_TOKEN set, the download fails with HTTP 401.

Running it on your laptop for a class demo. If one person is using the model, Ollama starts faster, installs in a minute and works on Windows. Use the right tool.

Try it yourself

Run fit.py after changing GPU_GIB to 24 — an RTX 3090 or 4090. Then look up config.json for a model you actually want to serve, read the four numbers out of it, and add it to the MODELS dictionary. That habit, checked before you rent anything, will save you more money than any other five minutes in this section.

What to learn next

  • Model deployment — where to put this server without going broke.
  • Ollama — the single-user version, for your own machine.
  • Context windows — why that number drives every memory decision here.

Researcher — Mathematics and papers.

PagedAttention

Kwon et al. (2023) identified that pre-vLLM serving systems allocated the KV cache as one contiguous tensor per sequence, sized to the maximum generation length. They measured 60–80% of that memory as waste, in three forms: internal fragmentation (reserved but never generated), external fragmentation (unusable gaps between allocations), and reservation (held for future tokens of a live sequence).

PagedAttention borrows virtual memory paging. The KV cache is partitioned into fixed-size blocks (typically 16 tokens). A per-sequence block table maps logical token positions to arbitrary physical blocks, so a sequence's cache need not be contiguous. The attention kernel is rewritten to gather across blocks.

Consequences:

  • Waste is bounded by one partially-filled block per sequence, so under 4% instead of 60–80%.
  • Copy-on-write block sharing becomes possible. Parallel samples from one prompt, or beam search branches, share prompt blocks with a reference count, and fork only on divergence.
  • The reported end-to-end result is 2–4× throughput at equal latency versus HuggingFace TGI and FasterTransformer of the time.

Continuous batching

Also called iteration-level scheduling; introduced as Orca (Yu et al., 2022). The scheduler runs per decoding step, not per request. After each forward pass, finished sequences are evicted and queued sequences admitted.

Static batching wastes capacity proportional to the variance in output length. Since output lengths are long-tailed in practice, that variance is large, and this is where most of the gain comes from.

Preemption matters at saturation. When the block pool is exhausted, vLLM either swaps blocks to host memory or recomputes them on resumption. Recomputation is usually cheaper, because prefill is compute-bound and PCIe transfer is not.

Prefill and decode are different workloads

PhaseParallelismBound byArithmetic intensity
Prefillall prompt tokens at oncecompute (FLOPs)high
Decodeone token per sequencememory bandwidth~1 op per weight byte

Mixing them in one batch means a long prefill blocks every decode step in flight, producing inter-token latency spikes. Chunked prefill (Agrawal et al., Sarathi, 2023) splits a long prefill into token-budgeted chunks and co-schedules each chunk with ongoing decodes, trading a small prefill slowdown for a large tail-latency improvement. It is on by default in current vLLM.

The alternative is disaggregated serving (Zhong et al., DistServe, 2024): run prefill and decode on separate GPU pools and ship the KV cache between them. This wins when the two phases have genuinely different SLOs, at the cost of an interconnect dependency.

Prefix caching

Requests that share a prefix — a long system prompt, a few-shot block, a shared document — can share KV blocks. vLLM hashes block contents and reuses matching blocks across requests (--enable-prefix-caching, default on in V1).

For a workload with a 2,000-token shared system prompt and 100-token user turns, this removes roughly 95% of prefill compute. It is frequently the largest single win available, and it costs one flag.

Capacity, via Little's law

L  =  λ · W

L is mean concurrency, λ the arrival rate, W the mean time in system. Concurrency is capped by the block pool:

L_max  =  total_kv_bytes / (kv_bytes_per_token × mean_sequence_tokens)

Substituting gives sustainable throughput. The practical reading: halving kv_bytes_per_token — through GQA, FP8 KV cache, or a smaller --max-model-len — doubles concurrency at fixed latency. This is why grouped-query attention (Ainslie et al., 2023) changed serving economics more than any kernel optimisation.

Note the failure mode when λ exceeds capacity: queueing delay grows without bound, and p99 latency degrades far faster than mean latency. Provision on p99, and measure with a load generator, not with time curl.

Parallelism

  • Tensor parallelism (--tensor-parallel-size) shards each weight matrix across GPUs. Requires an all-reduce per layer, so it wants NVLink; over PCIe the collective cost dominates quickly.
  • Pipeline parallelism (--pipeline-parallel-size) shards by layer. Cheaper interconnect, but introduces bubbles unless microbatched.
  • Expert parallelism for mixture-of-experts models, routing tokens to expert shards.

Rule of thumb: use the smallest tensor-parallel degree that fits the weights plus a workable KV budget. Sharding beyond necessity costs latency for nothing.

Quantization in serving

  • FP8 (E4M3) on Hopper and later — near-lossless, with hardware-accelerated matmul. The default choice where available.
  • AWQ (Lin et al., 2023) and GPTQ (Frantar et al., 2022) — 4-bit weight-only. Weights shrink 4×; activations stay 16-bit. Excellent for memory-bound decode, less useful for compute-bound prefill.
  • FP8 KV cache (--kv-cache-dtype fp8) — halves the cache, roughly doubling concurrency. Usually a better return than quantizing weights further.

Papers

What to learn next

  • Model deployment — where to put this server without going broke.
  • Ollama — the single-user version, for your own machine.
  • Context windows — why that number drives every memory decision here.

What to learn next

These follow on from what you just read.

  • LLM Development

    LangChain

    LangChain is a toolkit that joins prompts, models and parsers into one pipeline, so you can swap any part without rewriting the rest.

  • LLM Development

    LlamaIndex

    LlamaIndex loads your documents, cuts them into searchable pieces and finds the right piece when a question arrives, so a model can answer from your files.

  • LLM Development

    Function calling and tools

    Function calling lets a model ask your program to run a specific function with specific arguments, so it can use a calculator, a database or an API instead of guessing.