Text Generation Inference (TGI)
TGI is a server built specifically for generating text from large language models fast, using continuous batching to keep a GPU busy across requests of very different lengths.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
TGI is a server built specifically to generate text from large language models quickly, on real GPU hardware.
The analogy you have already lived
Imagine a line cook who takes new orders only after every current dish is fully plated. One order might be a quick sandwich, another a slow-roasting dish. The sandwich customer waits behind the roast, even though their order was ready to start immediately.
A smarter kitchen lets a new quick order start the moment a burner frees up. It does not wait for a slow dish already cooking to finish first. TGI, Hugging Face's serving engine for language models, is built around exactly that idea for text generation.
Why it exists
Generating text happens one word-piece at a time — a token, a small chunk of a word. Different requests need wildly different numbers of tokens: a one-line answer versus a long explanation. A server that waits for the slowest request in a batch before starting a new one wastes enormous GPU time.
TGI, along with similar engines like vLLM, was built specifically to solve this. New requests join generation already in progress. Finished ones leave, without the whole batch waiting on the slowest member.
How it works
naive batching: wait for EVERY request in the batch to finish
-> GPU sits idle waiting on the slowest one
continuous batching: as soon as ANY request finishes,
immediately slot a new one into its place
-> GPU stays busy, request by requestThis one idea, continuous batching, explains most of the gap. TGI, vLLM and similar engines serve far more users per GPU than a naive implementation.
A real example you have seen
A chatbot that starts showing you words as they are generated. It does this while simultaneously serving other users' very different-length questions. Nobody waits for the longest answer in the group to finish first.
The honest part
TGI runs as a specialised Rust-and-Python server, usually inside Docker. It expects real GPU hardware, with the model's weights downloaded — genuine infrastructure this lesson cannot spin up for you. What follows teaches the exact contract you talk to it through, marking what is a stand-in every time.
Remember this
- TGI is a specialised server for generating text from language models, not a general web framework.
- Continuous batching is its core trick — new requests join as soon as a slot frees up.
- It genuinely needs a GPU and real infrastructure to run — this lesson teaches the request contract honestly.
What to learn next
- Streaming token responses — the client-facing half of what a text-generation server produces, one token at a time.
- vLLM — TGI's closest peer, built around the same continuous-batching idea with its own scheduling design.
- KV-cache quantisation — shrinking the exact memory structure PagedAttention is managing, for even more concurrent requests.
Developer — Code and libraries.
Setup
pip install flask requestsReal TGI is launched with Docker, pointing at a Hugging Face model name:
docker run --gpus all -p 8080:80 \
ghcr.io/huggingface/text-generation-inference:latest \
--model-id mistralai/Mistral-7B-Instruct-v0.2Talking to TGI, without a GPU
TGI exposes a documented HTTP JSON contract at /generate. The Flask server below is not TGI — it is a small stand-in that speaks the same request and response shape, so the client code is what stays unchanged when pointed at a real TGI container.
import logging
import threading
import time
import requests
from flask import Flask, jsonify, request
from werkzeug.serving import make_server
logging.getLogger("werkzeug").setLevel(logging.ERROR)
# STAND-IN ONLY: real TGI does continuous batching, tensor-parallel GPU
# inference and quantised weight loading in a Rust process. This exists
# only to let TGI's real /generate wire contract run without a GPU.
FAKE_VOCAB = ["the", "loan", "is", "approved", "for", "fifty", "thousand", "rupees"]
app = Flask(__name__)
app.json.sort_keys = False # keep dict key order as written, not alphabetised
@app.post("/generate")
def generate():
body = request.get_json()
max_new_tokens = body.get("parameters", {}).get("max_new_tokens", 5)
words = (FAKE_VOCAB * ((max_new_tokens // len(FAKE_VOCAB)) + 1))[:max_new_tokens]
tokens = [{"id": i, "text": w, "logprob": -0.1 * (i + 1), "special": False}
for i, w in enumerate(words)]
return jsonify({
"generated_text": " ".join(words),
"details": {
"finish_reason": "length",
"generated_tokens": len(words),
"tokens": tokens,
},
})
server = make_server("127.0.0.1", 8973, app) # quieter than app.run(): no startup banner
threading.Thread(target=server.serve_forever, daemon=True).start()
time.sleep(0.6)
# ---- real TGI client call, unmodified from the real API ----
payload = {
"inputs": "Summarise the loan decision:",
"parameters": {"max_new_tokens": 6, "temperature": 0.7, "details": True},
}
r = requests.post("http://127.0.0.1:8973/generate", json=payload)
data = r.json()
print("generated_text:", data["generated_text"])
print("finish_reason :", data["details"]["finish_reason"])
print("first token :", data["details"]["tokens"][0])generated_text: the loan is approved for fifty
finish_reason : length
first token : {'id': 0, 'text': 'the', 'logprob': -0.1, 'special': False}Line-by-line walkthrough
parameters.max_new_tokens and details.tokens are TGI's real, documented field names — this is the exact shape a real TGI server returns, with a logprob (how confident the model was in that specific token) attached to each one.
finish_reason tells you why generation stopped: "length" means it hit max_new_tokens; a real server also returns "eos_token" when the model generated a natural stopping token on its own, or "stop_sequence" when a caller-provided stop string was matched.
Common mistakes
Setting max_new_tokens far higher than needed, as a precaution. Every unused token slot is GPU memory and time not spent on someone else's request — continuous batching helps, but it cannot recover capacity reserved and then unused.
Ignoring finish_reason. A response that stopped because it hit the length limit is often mid-sentence, not actually finished — treating it identically to a natural completion produces confusing, truncated output shown to users.
Assuming TGI and vLLM are interchangeable with zero testing. Both implement continuous batching, but with different scheduling details, quantisation support and API surfaces. Benchmark on your own workload before choosing.
Try it yourself
Extend the stand-in to also implement /generate_stream, returning tokens one at a time as Server-Sent Events instead of all at once. Streaming token responses covers exactly this pattern in general, against a plain FastAPI server.
What to learn next
- Streaming token responses — the client-facing half of what a text-generation server produces, one token at a time.
- vLLM — TGI's closest peer, built around the same continuous-batching idea with its own scheduling design.
- KV-cache quantisation — shrinking the exact memory structure PagedAttention is managing, for even more concurrent requests.
Researcher — Mathematics and papers.
Continuous batching, formally
Static batching groups $B$ requests and runs decoding steps until the longest sequence in the batch reaches its stop condition, wasting compute on every already-finished sequence in the batch during that tail. Continuous batching (iteration-level scheduling) instead re-evaluates the batch composition at every decoding step: a finished sequence is evicted and a new queued request admitted into its slot immediately, keeping the batch's occupied slots close to full throughout.
Orca (Yu et al., OSDI 2022) introduced this scheduling approach and reported substantial throughput improvements over request-level static batching on real generation workloads with heterogeneous output lengths — the common case in production, where request lengths are rarely known or uniform in advance.
Memory management: the KV cache problem
Autoregressive generation caches each previous token's key and value tensors (the KV cache) to avoid recomputing them at every step. This cache grows linearly with sequence length and must be allocated per request, and its size is not known in advance for a request whose eventual output length is unknown at admission time.
Naive contiguous allocation for the KV cache leads to severe memory fragmentation once requests of varying lengths are admitted and evicted continuously. vLLM's PagedAttention (Kwon et al., SOSP 2023) addresses this directly, borrowing virtual-memory paging: the KV cache is managed in fixed-size non-contiguous blocks, referenced by a per-sequence block table, eliminating fragmentation and allowing near-full GPU memory utilisation for cache. TGI has adopted comparable block-based cache management in more recent releases, converging with vLLM on the same underlying strategy.
Quantisation and throughput
TGI supports serving quantised weights (GPTQ, AWQ, and others — see GPTQ) directly, trading a measured, small accuracy loss for substantially reduced memory footprint per model instance, which in turn allows more concurrent requests or larger batch sizes within the same GPU memory budget. This interacts directly with continuous batching's throughput: a smaller per-instance memory footprint raises the practical ceiling on how many sequences can be resident and batched together at once.
Papers and systems
- Yu et al., Orca: A Distributed Serving System for Transformer-Based Generative Models, OSDI 2022 — the paper introducing continuous batching.
- Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM), SOSP 2023.
- Hugging Face's TGI documentation and source — the authoritative reference for its exact API contract and configuration: github.com/huggingface/text-generation-inference
What to learn next
- Streaming token responses — the client-facing half of what a text-generation server produces, one token at a time.
- vLLM — TGI's closest peer, built around the same continuous-batching idea with its own scheduling design.
- KV-cache quantisation — shrinking the exact memory structure PagedAttention is managing, for even more concurrent requests.