Serving Models in Production

Text Generation Inference (TGI)

TGI is a server built specifically for generating text from large language models fast, using continuous batching to keep a GPU busy across requests of very different lengths.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

TGI is a server built specifically to generate text from large language models quickly, on real GPU hardware.

The analogy you have already lived

Imagine a line cook who takes new orders only after every current dish is fully plated. One order might be a quick sandwich, another a slow-roasting dish. The sandwich customer waits behind the roast, even though their order was ready to start immediately.

A smarter kitchen lets a new quick order start the moment a burner frees up. It does not wait for a slow dish already cooking to finish first. TGI, Hugging Face's serving engine for language models, is built around exactly that idea for text generation.

Why it exists

Generating text happens one word-piece at a time — a token, a small chunk of a word. Different requests need wildly different numbers of tokens: a one-line answer versus a long explanation. A server that waits for the slowest request in a batch before starting a new one wastes enormous GPU time.

TGI, along with similar engines like vLLM, was built specifically to solve this. New requests join generation already in progress. Finished ones leave, without the whole batch waiting on the slowest member.

How it works

naive batching: wait for EVERY request in the batch to finish
                 -> GPU sits idle waiting on the slowest one

continuous batching: as soon as ANY request finishes,
                      immediately slot a new one into its place
                 -> GPU stays busy, request by request

This one idea, continuous batching, explains most of the gap. TGI, vLLM and similar engines serve far more users per GPU than a naive implementation.

A real example you have seen

A chatbot that starts showing you words as they are generated. It does this while simultaneously serving other users' very different-length questions. Nobody waits for the longest answer in the group to finish first.

The honest part

TGI runs as a specialised Rust-and-Python server, usually inside Docker. It expects real GPU hardware, with the model's weights downloaded — genuine infrastructure this lesson cannot spin up for you. What follows teaches the exact contract you talk to it through, marking what is a stand-in every time.

Remember this

  • TGI is a specialised server for generating text from language models, not a general web framework.
  • Continuous batching is its core trick — new requests join as soon as a slot frees up.
  • It genuinely needs a GPU and real infrastructure to run — this lesson teaches the request contract honestly.

What to learn next

  • Streaming token responses — the client-facing half of what a text-generation server produces, one token at a time.
  • vLLM — TGI's closest peer, built around the same continuous-batching idea with its own scheduling design.
  • KV-cache quantisation — shrinking the exact memory structure PagedAttention is managing, for even more concurrent requests.

Developer — Code and libraries.

Setup

bash
pip install flask requests

Real TGI is launched with Docker, pointing at a Hugging Face model name:

bash
docker run --gpus all -p 8080:80 \
  ghcr.io/huggingface/text-generation-inference:latest \
  --model-id mistralai/Mistral-7B-Instruct-v0.2

Talking to TGI, without a GPU

TGI exposes a documented HTTP JSON contract at /generate. The Flask server below is not TGI — it is a small stand-in that speaks the same request and response shape, so the client code is what stays unchanged when pointed at a real TGI container.

tgi_client_demo.py
import logging
import threading
import time

import requests
from flask import Flask, jsonify, request
from werkzeug.serving import make_server

logging.getLogger("werkzeug").setLevel(logging.ERROR)

# STAND-IN ONLY: real TGI does continuous batching, tensor-parallel GPU
# inference and quantised weight loading in a Rust process. This exists
# only to let TGI's real /generate wire contract run without a GPU.
FAKE_VOCAB = ["the", "loan", "is", "approved", "for", "fifty", "thousand", "rupees"]

app = Flask(__name__)
app.json.sort_keys = False  # keep dict key order as written, not alphabetised

@app.post("/generate")
def generate():
    body = request.get_json()
    max_new_tokens = body.get("parameters", {}).get("max_new_tokens", 5)
    words = (FAKE_VOCAB * ((max_new_tokens // len(FAKE_VOCAB)) + 1))[:max_new_tokens]
    tokens = [{"id": i, "text": w, "logprob": -0.1 * (i + 1), "special": False}
              for i, w in enumerate(words)]
    return jsonify({
        "generated_text": " ".join(words),
        "details": {
            "finish_reason": "length",
            "generated_tokens": len(words),
            "tokens": tokens,
        },
    })

server = make_server("127.0.0.1", 8973, app)  # quieter than app.run(): no startup banner
threading.Thread(target=server.serve_forever, daemon=True).start()
time.sleep(0.6)

# ---- real TGI client call, unmodified from the real API ----
payload = {
    "inputs": "Summarise the loan decision:",
    "parameters": {"max_new_tokens": 6, "temperature": 0.7, "details": True},
}
r = requests.post("http://127.0.0.1:8973/generate", json=payload)
data = r.json()
print("generated_text:", data["generated_text"])
print("finish_reason :", data["details"]["finish_reason"])
print("first token   :", data["details"]["tokens"][0])
Output
generated_text: the loan is approved for fifty
finish_reason : length
first token   : {'id': 0, 'text': 'the', 'logprob': -0.1, 'special': False}

Line-by-line walkthrough

parameters.max_new_tokens and details.tokens are TGI's real, documented field names — this is the exact shape a real TGI server returns, with a logprob (how confident the model was in that specific token) attached to each one.

finish_reason tells you why generation stopped: "length" means it hit max_new_tokens; a real server also returns "eos_token" when the model generated a natural stopping token on its own, or "stop_sequence" when a caller-provided stop string was matched.

Common mistakes

Setting max_new_tokens far higher than needed, as a precaution. Every unused token slot is GPU memory and time not spent on someone else's request — continuous batching helps, but it cannot recover capacity reserved and then unused.

Ignoring finish_reason. A response that stopped because it hit the length limit is often mid-sentence, not actually finished — treating it identically to a natural completion produces confusing, truncated output shown to users.

Assuming TGI and vLLM are interchangeable with zero testing. Both implement continuous batching, but with different scheduling details, quantisation support and API surfaces. Benchmark on your own workload before choosing.

Try it yourself

Extend the stand-in to also implement /generate_stream, returning tokens one at a time as Server-Sent Events instead of all at once. Streaming token responses covers exactly this pattern in general, against a plain FastAPI server.

What to learn next

  • Streaming token responses — the client-facing half of what a text-generation server produces, one token at a time.
  • vLLM — TGI's closest peer, built around the same continuous-batching idea with its own scheduling design.
  • KV-cache quantisation — shrinking the exact memory structure PagedAttention is managing, for even more concurrent requests.

Researcher — Mathematics and papers.

Continuous batching, formally

Static batching groups $B$ requests and runs decoding steps until the longest sequence in the batch reaches its stop condition, wasting compute on every already-finished sequence in the batch during that tail. Continuous batching (iteration-level scheduling) instead re-evaluates the batch composition at every decoding step: a finished sequence is evicted and a new queued request admitted into its slot immediately, keeping the batch's occupied slots close to full throughout.

Orca (Yu et al., OSDI 2022) introduced this scheduling approach and reported substantial throughput improvements over request-level static batching on real generation workloads with heterogeneous output lengths — the common case in production, where request lengths are rarely known or uniform in advance.

Memory management: the KV cache problem

Autoregressive generation caches each previous token's key and value tensors (the KV cache) to avoid recomputing them at every step. This cache grows linearly with sequence length and must be allocated per request, and its size is not known in advance for a request whose eventual output length is unknown at admission time.

Naive contiguous allocation for the KV cache leads to severe memory fragmentation once requests of varying lengths are admitted and evicted continuously. vLLM's PagedAttention (Kwon et al., SOSP 2023) addresses this directly, borrowing virtual-memory paging: the KV cache is managed in fixed-size non-contiguous blocks, referenced by a per-sequence block table, eliminating fragmentation and allowing near-full GPU memory utilisation for cache. TGI has adopted comparable block-based cache management in more recent releases, converging with vLLM on the same underlying strategy.

Quantisation and throughput

TGI supports serving quantised weights (GPTQ, AWQ, and others — see GPTQ) directly, trading a measured, small accuracy loss for substantially reduced memory footprint per model instance, which in turn allows more concurrent requests or larger batch sizes within the same GPU memory budget. This interacts directly with continuous batching's throughput: a smaller per-instance memory footprint raises the practical ceiling on how many sequences can be resident and batched together at once.

Papers and systems

  • Yu et al., Orca: A Distributed Serving System for Transformer-Based Generative Models, OSDI 2022 — the paper introducing continuous batching.
  • Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM), SOSP 2023.
  • Hugging Face's TGI documentation and source — the authoritative reference for its exact API contract and configuration: github.com/huggingface/text-generation-inference

What to learn next

  • Streaming token responses — the client-facing half of what a text-generation server produces, one token at a time.
  • vLLM — TGI's closest peer, built around the same continuous-batching idea with its own scheduling design.
  • KV-cache quantisation — shrinking the exact memory structure PagedAttention is managing, for even more concurrent requests.

What to learn next

These follow on from what you just read.

  • Serving Models in Production

    gRPC vs REST for inference

    gRPC and REST are two different ways for a caller to talk to a model server, trading gRPC's speed and strict structure against REST's simplicity and universal support.

  • Serving Models in Production

    Warm-up and cold starts

    A cold start is the extra time a fresh server instance needs before it answers at full speed, from starting the process through loading weights to its first few slow calls.

  • Serving Models in Production

    Serving many models on one machine

    Multi-model serving keeps several trained models loaded in one process, routing each request to the right one, instead of running a separate server per model.