Serving Models in Production

Serving many models on one machine

Multi-model serving keeps several trained models loaded in one process, routing each request to the right one, instead of running a separate server per model.

On this page 7
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Multi-model serving keeps several trained models loaded in one running process, routing each request to the right one.

The analogy you have already lived

A big Indian household kitchen often has several jars on one shelf, not a separate shelf per spice. One shelf, labelled jars, and whoever is cooking reaches for the one they need. Nobody builds an entire new shelf every time a new spice enters the house.

Serving many small models can work the same way. One process, several loaded models, routed by name. It avoids running a whole separate server for each one.

Why it exists

A real product often needs more than one model: a fraud model, a churn model, an upsell model, each trained separately. Running a full, separate server for each one adds up fast when models are individually small. Each one carries its own process, its own warm-up, its own memory overhead.

Keeping several small models in one process is often cheaper to operate than one server per model. Route each incoming request to the right one. This works as long as no single model is so large, or so heavily used, that it needs dedicated hardware to itself.

How it works

one running process
      |
      +---- risk-model    (loaded once at startup)
      +---- churn-model   (loaded once at startup)
      +---- upsell-model  (loaded once at startup)

request arrives:  "/predict/churn-model"
      |
      v
  look up "churn-model" in the loaded set, use it

The request itself names which model it wants. The server's only new job, beyond what model serving already covers, is picking the right one from several loaded options.

A real example you have seen

A bank's mobile app backend might score a transaction for fraud, check loan eligibility, and rank offers to show you. That is often three separate models, frequently served from shared infrastructure rather than three entirely separate systems.

Remember this

  • Multi-model serving keeps several models loaded in one process, routed by name.
  • It saves the overhead of running a full separate server per model, when each model is individually small.
  • The request has to say which model it wants — routing by name is the one genuinely new piece.

What to learn next

  • Ray Serve — a system built to manage exactly this kind of multi-model resource allocation across many machines.
  • Warm-up and cold starts — the cost paid every time a model has to be loaded, dynamically or otherwise.
  • Versioning and retiring features — the same registry discipline, worth applying to which model version each name currently points at.

Developer — Code and libraries.

Setup

bash
pip install fastapi "uvicorn[standard]" scikit-learn numpy httpx

Three models, one process, routed by name

multi_model_server.py
import numpy as np
from fastapi import FastAPI, HTTPException
from fastapi.testclient import TestClient
from pydantic import BaseModel
from sklearn.linear_model import LogisticRegression

def make_model(seed):
    rng = np.random.RandomState(seed)
    X = rng.uniform(0, 10, (200, 2))
    y = (X[:, 0] + X[:, 1] > 10).astype(int)
    return LogisticRegression().fit(X, y)

# Three small models loaded once at startup, kept in one dictionary.
# Each is a few KB here; in a real deployment you would watch total
# process memory, since it is the sum of every model you keep loaded.
MODELS = {
    "risk-model": make_model(0),
    "churn-model": make_model(1),
    "upsell-model": make_model(2),
}

app = FastAPI()

class Row(BaseModel):
    a: float
    b: float

@app.post("/predict/{model_name}")
def predict(model_name: str, row: Row):
    model = MODELS.get(model_name)
    if model is None:
        raise HTTPException(status_code=404, detail=f"no model named '{model_name}'")
    prob = float(model.predict_proba([[row.a, row.b]])[0, 1])
    return {"model": model_name, "probability": round(prob, 4)}

client = TestClient(app)

for name in ["risk-model", "churn-model", "unknown-model"]:
    r = client.post(f"/predict/{name}", json={"a": 6.0, "b": 5.0})
    print(name, "->", r.status_code, r.json())
Output
risk-model -> 200 {'model': 'risk-model', 'probability': 0.9253}
churn-model -> 200 {'model': 'churn-model', 'probability': 0.9385}
unknown-model -> 404 {'detail': "no model named 'unknown-model'"}

Line-by-line walkthrough

MODELS is loaded once, at import time — the same "load once" discipline from model serving, applied to a whole dictionary of models instead of a single one.

The {model_name} path parameter is the routing mechanism. A missing model returns a clear 404 naming exactly what was requested and not found, rather than a confusing internal error.

Common mistakes

Loading every model, even ones rarely used, permanently into memory. Total memory is the sum of every loaded model. For many small models, this is fine; for even a few large ones, it stops being fine quickly — that is when Triton or Ray Serve's per-model resource management starts to earn its complexity.

No isolation between models. One model's bug (an unhandled exception, a memory leak) can take down every other model sharing its process. Decide deliberately whether that risk is acceptable for your set of models, or whether some belong in their own process.

Silent fallback to a default model on an unknown name. A typo in a model name should be loud — a clear 404, as above — never quietly answered by whichever model happens to be first in the dictionary.

No per-model versioning. "churn-model" alone does not say which trained version is running. Combine this with the registry idea from versioning and retiring features, applied to models instead of features.

Try it yourself

Add a GET /models endpoint that lists every currently loaded model name. Then add a fourth model, and confirm it appears without touching any of the existing routing code.

What to learn next

  • Ray Serve — a system built to manage exactly this kind of multi-model resource allocation across many machines.
  • Warm-up and cold starts — the cost paid every time a model has to be loaded, dynamically or otherwise.
  • Versioning and retiring features — the same registry discipline, worth applying to which model version each name currently points at.

Researcher — Mathematics and papers.

The isolation-versus-efficiency trade-off

Co-locating models in one process trades fault and resource isolation for lower fixed overhead per model. The two extremes:

  • One process per model — full isolation (a crash in one cannot affect another), but pays a full process's worth of memory, startup time, and (for GPU models) device-memory overhead per model, regardless of how small or lightly used any individual model is.
  • All models in one process — minimal fixed overhead, but a shared fate: one model's memory leak, unhandled exception in a shared dependency, or resource exhaustion can degrade or take down every model sharing that process.

Production systems usually land between these extremes: grouping models by resource profile (all small, lightly used models together; large or heavily used models given dedicated capacity), rather than choosing one policy universally.

GPU memory as the harder constraint

CPU memory failures degrade gracefully in most cases (the OS pages, or the process is killed and restarted); GPU memory is a hard, fixed budget with no virtual memory equivalent in most inference frameworks. Loading $k$ models onto one GPU requires $\sum_i \text{mem}(m_i) \le \text{GPU memory}$, with no room for a model that does not fit, however lightly it might be used. This is precisely the resource-scheduling problem Triton's instance_group and model-loading policies, and Ray Serve's per-deployment resource requests, exist to manage explicitly rather than leaving to chance.

Dynamic loading and LRU eviction

For a very large number of models, each individually lightly used (a common shape in per-customer or per-tenant model deployments), keeping every model resident permanently is often infeasible. A dynamic loading policy — load a model on first request, evict the least-recently-used model when a memory ceiling is reached — trades a cold-start penalty (see warm-up and cold starts) on a cache miss for bounded total memory use. The design mirrors any LRU cache, with the specific twist that a "cache miss" here costs a full model load, often tens to thousands of milliseconds, rather than a cheap recomputation.

Papers and systems

  • NVIDIA Triton's model management documentation covers both explicit and dynamic (on-demand) model loading policies in detail: docs.nvidia.com/deeplearning/triton-inference-server
  • Crankshaw et al., Clipper: A Low-Latency Online Prediction Serving System, NSDI 2017 — early treatment of serving heterogeneous models behind one system with shared resource management.
  • Narayanan et al., Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning Workloads, OSDI 2020 — cluster-level scheduling for exactly this kind of shared, multi-model resource contention.

What to learn next

  • Ray Serve — a system built to manage exactly this kind of multi-model resource allocation across many machines.
  • Warm-up and cold starts — the cost paid every time a model has to be loaded, dynamically or otherwise.
  • Versioning and retiring features — the same registry discipline, worth applying to which model version each name currently points at.

What to learn next

These follow on from what you just read.

  • Serving Models in Production

    Streaming token responses

    Streaming sends each piece of a generated answer to the caller the moment it is ready, so a user sees the first word almost immediately instead of waiting for the whole reply.

  • Serving Models in Production

    Health checks and readiness probes

    Liveness and readiness are two different questions a server must answer separately, so traffic-routing systems never send real requests to an instance that is alive but not yet able to serve.

  • Batching and Concurrency

    Dynamic batching

    Dynamic batching holds a few incoming requests for a short moment so one model call can serve all of them together, trading a little latency for a lot more throughput.