MLOps

Model serving

Model serving is putting your trained model behind an address other programs can call, so an app can send one row and get an answer back in milliseconds.

Read these first

On this page 10
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. The two shapes of serving
  5. How it works
  6. The one extra address you always add
  7. Where you have already used one
  8. The honest part
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Model serving means putting your trained model behind an address that other programs can send questions to.

The analogy you have already lived

You have stood at a counter and asked for something. A chai stall, a bank window, a pharmacy. You said what you wanted, waited a few seconds, and got it. You never went into the kitchen or the back office.

A model sitting in your notebook is a cook who only cooks when you are personally standing beside them. Serving builds the counter — a window with a bell that anyone can ring, all day, without you being there.

Why it exists

You trained a model. It scores well. Right now the only way to use it is to open a notebook. Run six cells in order, then type in a value by hand.

Nothing in the real world can do that.

The app on someone's phone cannot open your notebook. The website cannot run your cells. The nightly job cannot type a value. They all speak one language: send a small message to an address, get a small message back.

Serving turns your model into that address.

The two shapes of serving

Online — one question at a time, answered right now. Someone taps "Pay". You have about fifty milliseconds. A fraud model has to answer inside that.

Batch — a huge pile of questions, answered overnight. Score every customer for next week's offer. Nobody is waiting, so it can take two hours.

Batch is far easier and often the right answer. Many teams build a complicated live service when a scheduled job at 2 a.m. would have done. Ask which one you actually need before you build either.

How it works

   phone app
      |
      |  "income 60, 7 years, age 34"   (a small message)
      v
  [ web server ]  <---- model file loaded ONCE when the server starts
      |
      |  checks the message is sensible
      |  puts the values in the order the model expects
      |  asks the model
      v
   "repaid: yes, 0.99"                  (a small message back)

Three details in that picture carry almost all the reliability.

Loaded once. The model is read from disk when the server starts, not when a request arrives. Reading it every time is the most common beginner mistake, and it gets worse as the model gets bigger.

Checked. Before the model sees anything, the server confirms the message has every field, with sensible values. An income of minus five is rejected at the door, not fed to the model.

In the right order. The server rebuilds the row using the column names the model was trained on. This is the failure from the first lesson, closed for good.

The one extra address you always add

A health check — a second address that answers "yes, I am alive and my model is loaded".

It sounds pointless. It is what the system in front of your server uses to decide whether to send traffic to it. Without one, a server that started but failed to load its model looks perfectly healthy and returns errors to real users.

Where you have already used one

  • UPI payments — a fraud score computed between your tap and the confirmation.
  • A shopping app's home page — ranked for you as the page loads.
  • Your keyboard's next-word suggestion — that one usually runs on the phone itself, which is serving too, without a network.
  • A support chat — every reply is a model call behind an address.

The honest part

Serving is where machine learning stops being a data problem and becomes a normal software problem.

Requests arrive at the same moment. Something times out. The server runs out of memory at 9 p.m. on a Friday. None of that is on any machine learning syllabus, and all of it will happen to you.

The good news: none of it is unique to machine learning. Web engineers have been solving it for thirty years, and their answers apply to you unchanged.

Remember this

  • Serving turns a model into an address other programs can call.
  • Load the model once at startup, and check every incoming message before the model sees it.
  • Batch scoring is easier than online scoring. Choose it whenever nobody is waiting.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install fastapi "uvicorn[standard]" scikit-learn pandas joblib httpx

httpx is what FastAPI's test client uses, so it is needed even though the example never opens a network connection.

Train something worth serving

train.py
import joblib
import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

rng = np.random.RandomState(0)
n = 400
X = pd.DataFrame({
    "income": rng.uniform(5, 80, n).round(1),     # thousands per month
    "years": rng.uniform(0, 10, n).round(1),      # years of credit history
    "age": rng.randint(21, 65, n).astype(float),
})
score = 0.05 * X["income"] + 0.35 * X["years"] + 0.01 * X["age"] - 3.5
y = (score + rng.normal(0, 0.8, n) > 0).astype(int)

# The scaler travels inside the pipeline, so serving cannot forget to apply it.
model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000)).fit(X, y)
joblib.dump(model, "model.joblib")
print("train accuracy:", round(model.score(X, y), 4), "  positives:", int(y.sum()), "/", n)
Output
train accuracy: 0.885   positives: 264 / 400

The make_pipeline line is the important one. Scaling lives inside the saved object, so the serving code physically cannot forget to apply it. A scaler that lives in a separate notebook cell is a production incident waiting for a date.

The service

serve.py
from contextlib import asynccontextmanager

import joblib
import pandas as pd
from fastapi import FastAPI
from pydantic import BaseModel, Field

STATE = {}


@asynccontextmanager
async def lifespan(app: FastAPI):
    # Loaded once when the process starts, not once per request.
    STATE["model"] = joblib.load("model.joblib")
    STATE["columns"] = list(STATE["model"].feature_names_in_)
    yield
    STATE.clear()


app = FastAPI(title="loan-scorer", lifespan=lifespan)


class Applicant(BaseModel):
    income: float = Field(ge=0, le=1000, description="thousands per month")
    years: float = Field(ge=0, le=60)
    age: float = Field(ge=18, le=100)


@app.get("/healthz")
def healthz():
    return {"status": "ok", "columns": STATE.get("columns")}


@app.post("/predict")
def predict(applicant: Applicant):
    # Rebuilt as a named frame in the model's own column order.
    row = pd.DataFrame([applicant.model_dump()])[STATE["columns"]]
    probability = float(STATE["model"].predict_proba(row)[0, 1])
    return {"repaid": int(probability >= 0.5), "probability": round(probability, 4)}

Exercise it without starting a server

TestClient runs the whole application in-process. No port, no background terminal, no cleanup — which is also exactly how you would test this in CI.

check.py
from fastapi.testclient import TestClient

from serve import app

with TestClient(app) as client:
    print("health   ", client.get("/healthz").json())

    good = {"income": 60.0, "years": 7.0, "age": 34}
    print("good     ", client.post("/predict", json=good).json())

    poor = {"income": 8.0, "years": 0.5, "age": 22}
    print("thin file", client.post("/predict", json=poor).json())

    missing = {"income": 60.0, "years": 7.0}
    r = client.post("/predict", json=missing)
    print("missing  ", r.status_code, r.json()["detail"][0]["loc"], r.json()["detail"][0]["msg"])

    negative = {"income": -5.0, "years": 7.0, "age": 34}
    r = client.post("/predict", json=negative)
    print("negative ", r.status_code, r.json()["detail"][0]["msg"])
Output
health    {'status': 'ok', 'columns': ['income', 'years', 'age']}
good      {'repaid': 1, 'probability': 0.9941}
thin file {'repaid': 0, 'probability': 0.0037}
missing   422 ['body', 'age'] Field required
negative  422 Input should be greater than or equal to 0

The last two lines are the ones that matter

A field is missing, and the answer is 422 with the exact field name. An income is negative, and the answer is 422 with the exact rule that was broken.

The model was never called. Nobody has to debug a strange prediction, because no prediction was made. The caller gets a message they can act on, in the same second.

Compare that with the alternative: income=-5 reaches a StandardScaler, comes out as a large negative number, and produces a confident rejection. Same input, no error, wrong answer, discovered in three months. Validation at the edge is the cheapest reliability you will ever buy.

The with TestClient(app) as client: form matters too. Without the with, the lifespan function never runs, STATE stays empty, and every request fails with KeyError. That one has cost many people an afternoon.

Run it for real

bash
uvicorn serve:app --host 127.0.0.1 --port 8000

Then open http://127.0.0.1:8000/docs for an interactive page generated from your Applicant class, where you can send requests by hand. No output block here: the startup banner and the reload behaviour vary by version and by flags.

About loading the model once

The advice is universal. The size of the win is not, and it is worth measuring rather than repeating.

python
import time, joblib, pandas as pd
row = pd.DataFrame([{"income": 60.0, "years": 7.0, "age": 34.0}])

model = joblib.load("model.joblib")
t = time.perf_counter()
for _ in range(200):
    model.predict_proba(row)
once = (time.perf_counter() - t) / 200 * 1000

t = time.perf_counter()
for _ in range(200):
    joblib.load("model.joblib").predict_proba(row)
each = (time.perf_counter() - t) / 200 * 1000
print(f"load once: {once:.2f} ms   reload each time: {each:.2f} ms   ratio {each/once:.0f}x")
Output
load once: 0.41 ms   reload each time: 0.93 ms   ratio 2x

Those exact milliseconds are from one laptop and yours will differ — timing output always does. The honest reading is that for a tiny logistic regression the penalty is about double, which is survivable.

The reason the rule is stated so firmly is what happens as the model grows. A 500 MB transformer takes seconds to load, not microseconds. Reloading it per request turns a 20 ms service into a 3 s one.

Measure your own case instead of trusting either the rule or a benchmark you did not run.

Common mistakes

Loading the model at import time instead of in lifespan. It appears to work. Then your test suite imports the module, the file is missing, and collection fails before a single test runs. Lifespan keeps startup failures where they belong — at startup.

Returning only a label. Return the probability as well, and the model version that produced it. Without those two fields you cannot investigate a complaint, and you cannot build the drift monitoring in the next lesson.

One worker for a CPU-bound model. Python's global interpreter lock means one process handles one prediction at a time. Run uvicorn --workers 4 behind a process manager, and remember each worker holds its own full copy of the model in memory.

No timeout on the caller's side. If your service hangs, every caller hangs with it, and the failure spreads outwards. Callers should set a timeout and have an answer ready for when it fires.

Serving one request per row when the caller has a thousand. Add a /predict/batch endpoint that accepts a list. One predict call on a thousand-row frame is dramatically faster than a thousand calls on one row, because the per-call overhead is paid once.

No version in the response. When two versions run side by side during a rollout, "which model said this?" is unanswerable without it.

Try it yourself

Add a POST /predict/batch endpoint taking list[Applicant], build one DataFrame from all of them, and call predict_proba once.

Then time 500 applicants both ways. The gap is larger than most people guess, and it is free.

What to learn next

Researcher — Mathematics and papers.

Latency, throughput, and the relationship between them

These are separate quantities, and optimising one routinely damages the other.

  • Latency — time for one request, reported as a distribution. p50, p95, p99, p999. Never a mean; the mean of a long-tailed distribution describes nobody's experience.
  • Throughput — completed requests per second.

Little's law relates them through concurrency:

$$L = \lambda W$$

where $L$ is the mean number of requests in the system, $\lambda$ the arrival rate and $W$ the mean time in system. It holds for any stable system regardless of arrival or service distribution.

The consequence that matters operationally: as utilisation $\rho$ approaches 1, queueing delay grows as $\rho / (1 - \rho)$. At 90% utilisation, expected queue time is nine times the service time. Capacity plans targeting high utilisation are implicitly targeting terrible tail latency.

Tail latency compounds across a request fan-out. If a page issues $n$ independent backend calls and waits for all of them, the probability that at least one lands in the p99 tail is $1 - 0.99^{n}$ — about 63% at $n = 100$. Dean and Barroso (2013), The Tail at Scale, is the canonical treatment, and its mitigations (hedged requests, tied requests, micro-partitioning) apply directly to model serving.

Batching

Accelerators are throughput devices with fixed per-call overhead. Serving requests one at a time leaves most of the hardware idle.

Dynamic batching holds arriving requests for up to $t_{\max}$ milliseconds or until $B$ requests accumulate, then runs one forward pass. This trades a bounded latency increase for a large throughput increase. NVIDIA Triton, TorchServe and TF Serving all implement it; the tunables are the batch size and the queue delay budget.

Continuous batching matters for autoregressive generation, where requests have wildly different output lengths. Static batching wastes the whole batch slot until the longest sequence finishes. Orca (Yu et al., 2022) introduced iteration-level scheduling, admitting new requests at each decoding step; vLLM (Kwon et al., 2023) added PagedAttention, managing the KV cache in non-contiguous fixed-size blocks and eliminating the fragmentation that had capped batch sizes. Reported throughput gains over prior systems are in the range of 2–4×.

Getting the model faster

TechniqueTypical speedupCost
Graph export (ONNX, TorchScript, torch.compile)1.2–3×Export friction, dynamic-shape limits
INT8 post-training quantisation2–4×Small accuracy loss, needs calibration data
4-bit weight quantisation (LLMs)Memory-bound gainMeasurable quality loss, kernel support varies
Structured pruning1.5–2×Requires fine-tuning to recover accuracy
DistillationLarge, model-dependentA full training run

The ordering advice is unchanged for a decade: measure first. A large share of "slow model" reports are dominated by feature computation, JSON serialisation, or a database call in the request path — not by the forward pass.

Deployment strategies

  • Blue-green — two full environments, switch traffic atomically. Fast rollback, double the resources.
  • Canary — route a small percentage to the new version, watch metrics, increase gradually. Needs enough traffic for the comparison to have power.
  • Shadow — send a copy of live traffic to the new model, discard its output, compare offline. The only method that carries zero user risk. Costs double inference, and cannot measure anything downstream of the user seeing a result.
  • Interleaving — for ranking systems, mix results from both models in one list. Far more statistically efficient than an A/B test at equal traffic (Chapelle et al., 2012).

Shadow deployment deserves specific emphasis for ML because offline accuracy and online behaviour diverge routinely, and shadow traffic is the cheapest way to find out before users do.

Cold start and autoscaling

Scale-to-zero is attractive for cost and hostile to latency. The cold path is: schedule a node, pull the image, start the process, load weights, warm the caches, then serve. For a multi-gigabyte image and a large model this is minutes.

Mitigations in decreasing order of effect: keep a warm floor of instances, pre-pull images onto nodes, load weights from a local NVMe cache rather than object storage, and scale on queue depth rather than CPU — CPU is a lagging indicator for a GPU-bound service and will consistently scale too late.

Papers and systems

  • Dean and Barroso, The Tail at Scale, CACM 2013 — research.google/pubs/pub40801
  • Crankshaw et al., Clipper: A Low-Latency Online Prediction Serving System, NSDI 2017 — arxiv.org/abs/1612.03079
  • Yu et al., Orca: A Distributed Serving System for Transformer-Based Generative Models, OSDI 2022
  • Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM), SOSP 2023 — arxiv.org/abs/2309.06180
  • Olston et al., TensorFlow-Serving: Flexible, High-Performance ML Serving, 2017 — arxiv.org/abs/1712.06139
  • Chapelle et al., Large-Scale Validation and Analysis of Interleaved Search Evaluation, TOIS 2012

What to learn next