Streaming token responses
Streaming sends each piece of a generated answer to the caller the moment it is ready, so a user sees the first word almost immediately instead of waiting for the whole reply.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Streaming sends each piece of a generated answer as soon as it is ready. The caller does not wait for the whole thing.
The analogy you have already lived
Watching a cricket match live, ball by ball, feels completely different from reading the final scorecard an hour later. The final score ends up the same either way. But seeing it unfold, piece by piece, feels immediate. Waiting for the complete result feels slow, even if the total wait is not actually that much longer.
Streaming a model's answer is choosing the live commentary, instead of the final scorecard.
Why it exists
A language model generates its answer one small piece at a time. Each piece is a token, roughly a word or part of a word. A non-streaming server waits until every token is generated, then sends the whole answer back in one go.
For a short answer, that wait is barely noticeable. For a long, detailed answer, a user can stare at a blank screen for several seconds before seeing anything at all. That is true even though the first few words were ready almost immediately.
How it works
non-streaming: generate ALL tokens -> send the WHOLE answer at once
user waits the FULL time before seeing anything
streaming: generate token -> send it immediately -> repeat
user sees the FIRST word almost instantly,
the rest arrives while they are already readingThe total time to the very last word is about the same either way. What changes dramatically is time to first word — how long the user stares at nothing before anything appears.
A real example you have seen
A chat assistant where the reply appears word by word as you watch. Compare that to the reply vanishing for several seconds, then dumping the entire answer on screen at once. That felt difference is exactly what streaming buys.
Remember this
- Streaming sends each generated piece as it becomes ready, not all at once at the end.
- The total generation time barely changes — what improves dramatically is time to first response.
- It matters most for long answers, where the wait for a full non-streamed reply would be most noticeable.
What to learn next
- Text Generation Inference (TGI) — the server-side scheduling that determines how smoothly a stream actually delivers under real load.
- gRPC vs REST for inference — a structurally different way to stream, worth understanding as an alternative to Server-Sent Events.
- Serving a model with FastAPI — the concurrency model a streaming endpoint has to share the server with.
Developer — Code and libraries.
Setup
pip install fastapi "uvicorn[standard]" requestsMeasuring time-to-first-token against a real running server
This starts a real Uvicorn server on a background thread, so the streaming behaviour measured is genuine network streaming, not an in-process shortcut.
import time
import threading
import asyncio
import requests
import uvicorn
from fastapi import FastAPI
from fastapi.responses import StreamingResponse
app = FastAPI()
WORDS = ["The", "loan", "is", "approved", "for", "an", "amount", "of",
"rupees", "fifty", "thousand", "."]
async def token_stream():
for word in WORDS:
await asyncio.sleep(0.05) # standing in for one decoding step
yield word + " "
@app.get("/generate-stream")
async def generate_stream():
return StreamingResponse(token_stream(), media_type="text/plain")
@app.get("/generate-full")
async def generate_full():
await asyncio.sleep(0.05 * len(WORDS)) # same total work, done all at once
return {"text": " ".join(WORDS)}
config = uvicorn.Config(app, host="127.0.0.1", port=8971, log_level="warning")
server = uvicorn.Server(config)
thread = threading.Thread(target=server.run, daemon=True)
thread.start()
while not server.started:
time.sleep(0.02)
t0 = time.perf_counter()
first_chunk_at = None
chunks = []
with requests.get("http://127.0.0.1:8971/generate-stream", stream=True) as r:
for chunk in r.iter_content(chunk_size=None, decode_unicode=True):
if first_chunk_at is None:
first_chunk_at = time.perf_counter() - t0
chunks.append(chunk)
stream_total = time.perf_counter() - t0
t0 = time.perf_counter()
r = requests.get("http://127.0.0.1:8971/generate-full")
full_total = time.perf_counter() - t0
print(f"streaming: time to FIRST token : {first_chunk_at*1000:.0f} ms")
print(f"streaming: time to LAST token : {stream_total*1000:.0f} ms")
print(f"non-streaming: time to full reply : {full_total*1000:.0f} ms")
print(f"streamed text: {''.join(chunks)!r}")
server.should_exit = True
thread.join(timeout=2)streaming: time to FIRST token : 58 ms streaming: time to LAST token : 629 ms non-streaming: time to full reply : 613 ms streamed text: 'The loan is approved for an amount of rupees fifty thousand . '
These are real timings from one run against a real local server — small variations between runs are normal. The result that matters is the shape: time to the last token is about the same either way, but time to the first token is roughly ten times faster when streaming. That gap is the entire point.
Line-by-line walkthrough
token_stream is an async generator — yield sends one chunk to the caller immediately, then the function pauses until the next yield, rather than building the whole response before returning anything.
StreamingResponse(token_stream(), ...) is what tells FastAPI to send each yielded chunk to the network as it is produced, instead of waiting to collect the full response first. This demo deliberately uses a real running server (uvicorn.Server, not the in-process test client from model serving) because true incremental delivery depends on the real network layer — an in-process test transport can buffer the whole response before handing it back, hiding the exact effect this lesson is trying to show.
Common mistakes
Testing streaming with an in-process test client and concluding it works. As noted above, some in-process transports buffer the full response before yielding anything, making a broken streaming endpoint look identical to a working one in that specific test setup. Verify against a real running server before trusting the result.
Streaming without a way for the client to tell chunks apart. Plain text streaming, as above, is simple but ambiguous for structured data. Real LLM APIs use Server-Sent Events (data: {...}\n\n per chunk) specifically so a client can parse each piece as a distinct JSON object.
Buffering on a proxy in front of the server. Some reverse proxies buffer responses by default, silently turning a streaming endpoint back into a non-streaming one before it reaches the user. This needs explicit proxy configuration to disable, and is a common reason "streaming works locally but not in production".
Forgetting the client has to actually consume the stream incrementally. A client library that reads response.text in one call defeats streaming entirely, even against a correctly streaming server — the client code has to iterate chunk by chunk, as r.iter_content does above.
Try it yourself
Change media_type="text/plain" to a proper Server-Sent Events format — yield f"data: {word}\n\n" instead of word + " ", with media_type="text/event-stream" — and update the client to parse data: lines. This is the format real LLM streaming APIs actually use.
What to learn next
- Text Generation Inference (TGI) — the server-side scheduling that determines how smoothly a stream actually delivers under real load.
- gRPC vs REST for inference — a structurally different way to stream, worth understanding as an alternative to Server-Sent Events.
- Serving a model with FastAPI — the concurrency model a streaming endpoint has to share the server with.
Researcher — Mathematics and papers.
Time-to-first-token as the primary streaming metric
For an interactive workload, time to first token (TTFT) and inter-token latency (the gap between consecutive tokens once generation is underway) are the two metrics that matter, not total generation time alone. A system can have identical total latency to a non-streaming system while dramatically improving perceived responsiveness, purely by front-loading delivery — exactly what the demo above measures.
TTFT is dominated by queueing delay (how long a request waits before its first decoding step runs) and prompt processing time (the forward pass over the input prompt, before any output token is produced). Inter-token latency is dominated by decoding step time, which scales with batch occupancy and, for continuously batched servers, by how many other sequences are sharing the same decoding step.
Streaming and continuous batching interact
A naive mental model treats streaming as a purely client-facing concern, independent of server-side batching. In practice they are coupled: continuous batching (see TGI) determines when each sequence's next token becomes available, which directly sets that sequence's realised inter-token latency for a streamed response. A server that batches well but streams poorly still delivers a good TTFT badly; the two need to be engineered together, not treated as separate concerns solved independently.
Backpressure in a streaming response
A slow client (a poor network connection, a client that stops reading) applies backpressure up the stack: if the server does not respect that backpressure, it can buffer an unbounded amount of generated-but-undelivered output in memory, holding GPU-resident generation state (the KV cache) for a request whose output the client may never fully consume. Production streaming servers bound this either by pausing generation when the write buffer is full, or by imposing a maximum buffered-output size per request and aborting generation past it.
Protocol choices for streaming
Server-Sent Events (a single long-lived HTTP response, chunked) is the dominant choice for LLM token streaming today, favoured for its simplicity over a WebSocket (a full bidirectional channel, more machinery than a one-directional token stream needs) or gRPC server-streaming (see gRPC vs REST for inference), which offers a more structurally native fit at the cost of needing a gRPC-aware client, unavailable directly in a browser.
Papers and systems
- Yu et al., Orca: A Distributed Serving System for Transformer-Based Generative Models, OSDI 2022 — iteration-level scheduling, the mechanism that determines per-token delivery timing under load.
- The OpenAI and Anthropic API documentation for streaming completions are the most widely adopted practical references for the SSE-based token-streaming contract in production use today.
- Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM), SOSP 2023 — the KV-cache backpressure concern above is a direct consequence of the memory model this paper addresses.
What to learn next
- Text Generation Inference (TGI) — the server-side scheduling that determines how smoothly a stream actually delivers under real load.
- gRPC vs REST for inference — a structurally different way to stream, worth understanding as an alternative to Server-Sent Events.
- Serving a model with FastAPI — the concurrency model a streaming endpoint has to share the server with.