Latency, Load Testing and Capacity
Time to first token vs tokens per second
How fast the first word of a reply appears and how fast the words keep coming afterward are two different numbers, and a chat product needs to care about both, not only one.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Time to first token is how long until the reply starts. Tokens per second is how fast it keeps going after that.
The analogy you have already lived
You have eaten at a restaurant that serves a starter quickly, then takes a while over the main course. Two different clocks matter to you: how long until something arrives, and how long the whole meal takes once it starts coming.
A model generating text has exactly the same two clocks. Time to first token is the wait for the starter. Tokens per second is the pace of everything that follows.
Why it exists
A language model does not produce a full reply in one instant. It reads your entire prompt first — a step called prefill — and only then starts writing, one token at a time. A token is a small chunk of text, often close to a word, that the model produces in a single step.
Time to first token, or TTFT, measures the wait from when you asked to when the very first token appears. It is dominated by prefill — a longer prompt takes longer to read before writing can even start.
Tokens per second, sometimes called decode speed, measures the pace once generation is underway — how quickly each following token arrives after the first one.
These two numbers can move independently. A system can start fast and then generate slowly, or start slowly and then generate quickly. Each combination feels different to a person watching the words appear.
How it works
you send a prompt
|
| PREFILL: reading the whole prompt <- this delay is
v time to first token
"The" <-- first token appears
|
| DECODE: one token at a time
v
"The" "answer" "is" "forty" "two" <-- pace here is
tokens per secondA real example you have seen
ChatGPT and similar chat apps show your reply appearing word by word instead of all at once. The delay before that first word is time to first token. How smoothly the rest of the words keep flowing afterward is tokens per second.
The honest part
Optimising one of these numbers can hurt the other. Making prefill faster for long prompts can compete for the same hardware resources that decode speed depends on. There is no single "speed" number for a language model — there are at least these two, and sometimes they trade off against each other.
Remember this
- Time to first token is the wait before generation starts, driven mostly by prefill.
- Tokens per second is the pace once generation is underway, driven by decode.
- A fast start and a fast pace are two separate engineering problems, not one.
What to learn next
- Continuous batching for LLMs — the scheduling technique that most directly improves tokens per second under real concurrent load.
- The KV cache — the mechanism that makes decode cheaper than reprocessing the whole conversation on every token.
- Capacity planning with Little's Law — reasoning about TTFT under real, queued traffic rather than in isolation.
Developer — Code and libraries.
Setup
pip install fastapi "uvicorn[standard]" httpxA real streaming server
Save this as stream_server.py and run it with uvicorn stream_server:app --host 127.0.0.1 --port 8942.
import time
from fastapi import FastAPI
from fastapi.responses import StreamingResponse
app = FastAPI()
PREFILL_TIME = 0.15 # reading the prompt, before the first word appears
PER_TOKEN_TIME = 0.02 # generating each following word
def token_stream(n_tokens: int):
time.sleep(PREFILL_TIME) # prefill: the model reads the whole prompt first
words = ("the quick brown fox jumps over the lazy dog again and again " * 3).split()
for i in range(n_tokens):
time.sleep(PER_TOKEN_TIME) # one decode step: one new word
yield words[i % len(words)] + " "
@app.get("/generate")
def generate(n_tokens: int = 30):
return StreamingResponse(token_stream(n_tokens), media_type="text/plain")Measuring both numbers over a real HTTP stream
import time
import httpx
URL = "http://127.0.0.1:8942/generate"
if __name__ == "__main__":
start = time.perf_counter()
first_token_at = None
n_words = 0
with httpx.stream("GET", URL, params={"n_tokens": 40}, timeout=30) as r:
for chunk in r.iter_text():
if first_token_at is None:
first_token_at = time.perf_counter()
n_words += len(chunk.split())
total = time.perf_counter() - start
ttft = first_token_at - start
tokens_per_sec = n_words / (total - ttft) if total > ttft else float("inf")
print(f"time to first token : {ttft*1000:6.1f} ms")
print(f"total time : {total*1000:6.1f} ms")
print(f"words received : {n_words}")
print(f"generation speed : {tokens_per_sec:5.1f} words/s after the first one")time to first token : 189.7 ms total time : 1007.0 ms words received : 40 generation speed : 48.9 words/s after the first one
Real measurements over a real, local HTTP stream. PREFILL_TIME is set to 150ms and PER_TOKEN_TIME to 20ms in the server, so the expected total is roughly 150 + 40*20 = 950ms — the measured 1007ms includes real network and processing overhead on top of that, which is genuine, not an error. Your own numbers will differ with machine speed and network stack, but time-to-first-token landing near 190ms against a 150ms prefill, and roughly 50 words/s against a 20ms-per-token server, should both hold.
Line-by-line walkthrough
StreamingResponse(token_stream(...)) sends each yielded chunk to the client as soon as it is produced, rather than waiting for the whole generator to finish — this is what makes streaming visible over the network at all.
httpx.stream(...) with r.iter_text() reads the response incrementally, chunk by chunk, as they arrive — recording first_token_at the instant the very first chunk shows up.
The tokens_per_sec calculation deliberately excludes the prefill wait, dividing only by total - ttft — mixing prefill time into a decode-speed number would understate it unfairly.
Common mistakes
Reporting one blended "requests per second" number for a streaming service. It hides both real numbers behind a single average that describes neither well. Report TTFT and tokens/sec separately.
Measuring TTFT with a short, unrealistic test prompt. Prefill time scales with prompt length. A TTFT measured with a five-word prompt tells you little about a service handling long documents or long conversation histories.
Blaming decode speed for a slow-feeling response, when prefill is the real cause. A long prompt with a genuinely fast decode speed can still feel slow to a user, entirely because of the prefill wait before anything appears.
Try it yourself
Change PREFILL_TIME from 0.15 to 1.0, simulating a much longer prompt, and rerun. Time to first token should rise by roughly the same amount — while generation speed barely changes, since decode was never the bottleneck here.
What to learn next
- Continuous batching for LLMs — the scheduling technique that most directly improves tokens per second under real concurrent load.
- The KV cache — the mechanism that makes decode cheaper than reprocessing the whole conversation on every token.
- Capacity planning with Little's Law — reasoning about TTFT under real, queued traffic rather than in isolation.
Researcher — Mathematics and papers.
Why prefill and decode have different cost profiles
Prefill processes the entire prompt in one forward pass, computing attention over every prompt token at once. It is largely compute-bound: for a prompt of length $n$, the dominant attention cost scales with $O(n^2)$ before any accounting for sequence-length optimisations (see quadratic attention cost and FlashAttention), and GPU utilisation during prefill is typically high, since there is a large batch of tokens to process together.
Decode produces one token per forward pass, reusing the KV cache so each step only computes attention for the new token against the cached ones. Decode is largely memory-bandwidth-bound: each step moves the model's full weights (and the growing KV cache) through memory to compute a comparatively small amount of new arithmetic, so GPU compute units sit relatively idle waiting on memory transfers. This is the standard justification for why batching (see continuous batching) helps decode throughput disproportionately — it amortises that memory traffic across more concurrent sequences.
TTFT as a function of prompt length and queueing
For a request arriving at a server already processing other requests, observed TTFT is prefill compute time plus however long the request waited for a scheduling slot — which is why TTFT under real load is better modelled with the queueing framework in capacity planning and Little's Law than treated as a fixed constant per prompt length.
Reporting conventions in the literature and industry benchmarks
Common reported metrics beyond the two in this lesson: inter-token latency (ITL), the time between consecutive tokens (the reciprocal of instantaneous tokens/sec, useful for spotting jitter within a single generation); end-to-end latency, TTFT plus total generation time for a fixed output length; and goodput, throughput measured only over requests that met a latency target, rather than raw throughput regardless of how badly individual requests missed their target (Kwon et al., 2023 discuss goodput specifically as the more decision-relevant metric for a serving system).
References
- Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM), SOSP 2023 — arxiv.org/abs/2309.06180
- Patel et al., Splitwise: Efficient Generative LLM Inference Using Phase Splitting, 2023 — proposes physically separating prefill and decode onto different hardware, since their resource profiles differ so much.
- NVIDIA, LLM Inference Benchmarking documentation — a practical reference for the standard metric definitions used across serving-system comparisons.
What to learn next
- Continuous batching for LLMs — the scheduling technique that most directly improves tokens per second under real concurrent load.
- The KV cache — the mechanism that makes decode cheaper than reprocessing the whole conversation on every token.
- Capacity planning with Little's Law — reasoning about TTFT under real, queued traffic rather than in isolation.