Serving Models in Production

NVIDIA Triton inference server

Triton is a dedicated inference server that handles dynamic batching, multiple frameworks and GPU scheduling for you, instead of you hand-rolling that logic inside a web framework.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Triton is a ready-made server built specifically for running trained models fast. You do not have to build that machinery yourself.

The analogy you have already lived

Serving a model with FastAPI is like being your own waiter, cook and cashier in a small stall you built yourself. It works, and you control everything.

Triton is more like a professional commercial kitchen built for exactly one job. It takes in orders and gets food out fast, at volume, with equipment tuned for that specific purpose. You still bring the recipe — the trained model. But the kitchen itself is not something you built by hand.

Why it exists

A hand-built FastAPI server sends one request through the model at a time, or needs custom code to batch several together. At real scale, with a GPU sitting mostly idle between small requests, that leaves a lot of speed on the table.

Triton was built by NVIDIA specifically to solve this. It batches incoming requests automatically, and runs several different model frameworks side by side. It squeezes far more use out of an expensive GPU than a general-purpose web server would, without you writing that logic yourself.

How it works

requests arrive, one by one, close together in time
      |
      v
  Triton HOLDS them briefly, collects several into one batch
      |
      v
  ONE forward pass through the model, for the whole batch
      |
      v
  Triton splits the answers back out, one per original request

Batching here means combining several requests into one bigger computation. A GPU is often faster doing one big job than the same work as many small ones.

A real example you have seen

A large e-commerce site's product recommendation service, where thousands of "what should I show this shopper" requests arrive every second. This is the kind of high-volume, GPU-backed traffic Triton is built for, rather than an occasional request from a small internal tool.

The honest part

Triton is real infrastructure. It runs as a specialised server process, usually inside Docker, often on a GPU machine you do not have on a laptop. This lesson explains exactly what it does and how to talk to it. Running the full server yourself needs that real infrastructure, which the demo below is honest about substituting.

Remember this

  • Triton is a purpose-built server for running trained models fast, not a general web framework.
  • Its main trick is automatic batching — combining nearby requests into one larger computation.
  • It genuinely needs real infrastructure (Docker and a GPU) to run for real — this lesson teaches the contract honestly.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install flask requests

Real Triton runs as a Docker container (docker run nvcr.io/nvidia/tritonserver) with a model repository — a folder per model containing the weights and a config.pbtxt file describing its inputs, outputs and batching settings. A minimal config.pbtxt looks like this:

config.pbtxt
name: "risk_model"
platform: "onnxruntime_onnx"
max_batch_size: 32
dynamic_batching {
  max_queue_delay_microseconds: 5000
}
input [
  { name: "INPUT0", data_type: TYPE_FP32, dims: [ 2 ] }
]
output [
  { name: "OUTPUT0", data_type: TYPE_FP32, dims: [ 1 ] }
]

dynamic_batching is the setting from the beginner section, made concrete: wait up to 5 milliseconds to collect a bigger batch before running.

Talking to Triton, without installing Triton

Triton speaks a documented, standard HTTP/JSON protocol (the KServe v2 inference protocol). The Flask server below is not Triton — it is a small stand-in that speaks the exact same wire contract, so the client code beneath it is unchanged if you point it at a real Triton container later.

triton_client_demo.py
import logging
import threading
import time

import numpy as np
import requests
from flask import Flask, jsonify, request
from sklearn.linear_model import LogisticRegression
from werkzeug.serving import make_server

logging.getLogger("werkzeug").setLevel(logging.ERROR)

rng = np.random.RandomState(0)
X = rng.uniform(0, 10, (200, 2))
y = (X[:, 0] + X[:, 1] > 10).astype(int)
model = LogisticRegression().fit(X, y)

# STAND-IN ONLY: real Triton does dynamic batching, multi-backend support
# (ONNX, TensorRT, PyTorch, Python) and GPU scheduling that this Flask app
# does not. It exists so the client below runs without a GPU or Docker.
app = Flask(__name__)

@app.get("/v2/health/ready")
def ready():
    return "", 200

@app.post("/v2/models/<model_name>/infer")
def infer(model_name):
    body = request.get_json()
    inp = body["inputs"][0]
    row = np.array(inp["data"]).reshape(inp["shape"])
    probs = [round(p, 4) for p in model.predict_proba(row)[:, 1].tolist()]
    return jsonify({
        "model_name": model_name,
        "model_version": "1",
        "outputs": [
            {"name": "OUTPUT0", "shape": [len(probs), 1], "datatype": "FP32", "data": probs}
        ],
    })

server = make_server("127.0.0.1", 8972, app)  # quieter than app.run(): no startup banner
threading.Thread(target=server.serve_forever, daemon=True).start()
time.sleep(0.6)

# ---- this is real Triton client code, unmodified by the stand-in above ----
health = requests.get("http://127.0.0.1:8972/v2/health/ready")
print("server ready:", health.status_code == 200)

payload = {
    "inputs": [
        {"name": "INPUT0", "shape": [1, 2], "datatype": "FP32", "data": [6.0, 5.0]}
    ]
}
r = requests.post("http://127.0.0.1:8972/v2/models/risk_model/infer", json=payload)
print("infer response:", r.json())
Output
server ready: True
infer response: {'model_name': 'risk_model', 'model_version': '1', 'outputs': [{'data': [0.9253], 'datatype': 'FP32', 'name': 'OUTPUT0', 'shape': [1, 1]}]}

Line-by-line walkthrough

The request and response shapes — inputs, shape, datatype, data, and the mirrored outputs — are Triton's real, documented protocol, not something invented for this lesson. That is the whole point of building the stand-in this way: this client code needs no change to work against a real Triton server, only a different port and hostname.

GET /v2/health/ready mirrors Triton's real readiness endpoint, the same idea covered in general terms in health checks and readiness probes.

Common mistakes

Assuming Triton is a drop-in replacement for a FastAPI server with no other changes. Triton expects models in specific formats (ONNX, TensorRT, TorchScript, or a custom Python backend) laid out in its model-repository structure — a plain joblib file will not load directly.

Setting max_queue_delay_microseconds too high. This is the batching wait time from the config file above. Set it too high and you trade away latency for a batching benefit that may not be worth it at your actual traffic level — measure, do not guess.

Reaching for Triton before you have GPU-bound traffic that justifies it. For light, CPU-only, low-volume traffic, a FastAPI server is simpler to operate and easier to debug. Triton earns its complexity at real GPU scale.

Try it yourself

Extend the stand-in's /v2/models/<model_name>/infer route to accept a batch of more than one row in data, and confirm the response returns one probability per row — that batch handling is the behaviour real Triton automates for you across many separate incoming requests.

What to learn next

Researcher — Mathematics and papers.

Dynamic batching, formally

Given requests arriving as a stream with inter-arrival distribution $A$, dynamic batching accumulates requests into a batch until either $B$ requests have arrived or $t_{\max}$ milliseconds have elapsed since the first request in the batch, whichever comes first. This bounds worst-case added latency to $t_{\max}$ while improving GPU utilisation, since a single forward pass over a batch of $B$ typically costs far less than $B$ separate forward passes on hardware with high per-call fixed overhead.

The throughput-latency trade-off is tuned by $B$ and $t_{\max}$ jointly: larger values improve throughput and worsen tail latency, following the same queueing-delay argument as model serving's researcher section — as the effective batch-collection utilisation approaches its ceiling, added wait time grows sharply, not linearly.

Model concurrency and multi-backend execution

Beyond batching a single model, Triton supports running multiple model instances concurrently on one GPU (instance_group in the config), and multiple different models — potentially in different frameworks — in the same server process, each with independent batching and instance settings. This directly enables multi-model serving at a scale a single hand-rolled process rarely reaches cleanly, since Triton manages per-model GPU memory and scheduling as first-class concerns rather than application code.

Backends

Triton's backend abstraction lets it execute ONNX Runtime, TensorRT, LibTorch, TensorFlow SavedModel, and arbitrary Python code (the Python backend) behind one uniform protocol. TensorRT backends specifically compile a model graph into a GPU-specific optimised execution plan ahead of time, trading a slow, hardware-specific build step for substantially lower per-inference latency than an unoptimised graph — directly relevant to warm-up and cold starts, since this compilation is part of what makes a fresh instance's first requests slow.

Papers and systems

  • NVIDIA Triton Inference Server documentation — the model repository format, config.pbtxt schema and KServe v2 protocol used verbatim above: docs.nvidia.com/deeplearning/triton-inference-server
  • The KServe v2 inference protocol specification — the standard this lesson's stand-in implements, shared across Triton, TorchServe and Seldon: kserve.github.io
  • Crankshaw et al., Clipper: A Low-Latency Online Prediction Serving System, NSDI 2017 — earlier academic work on the same dynamic-batching idea Triton implements in production.

What to learn next

  • BentoML — a Python-first alternative that trades some of Triton's raw throughput for much simpler local development.
  • Serving many models on one machine — the concurrency model Triton's multi-model support is built to handle.
  • Docker for ML — the packaging real Triton deployments are built on.

What to learn next

These follow on from what you just read.

  • Serving Models in Production

    BentoML

    BentoML is a Python framework that turns a trained model into a packaged, servable API with far less boilerplate than writing every route and Dockerfile by hand.

  • Serving Models in Production

    Ray Serve

    Ray Serve runs several copies of a model across many processes or machines behind one address, scaling a service by adding replicas instead of rewriting it.

  • Serving Models in Production

    Text Generation Inference (TGI)

    TGI is a server built specifically for generating text from large language models fast, using continuous batching to keep a GPU busy across requests of very different lengths.