Serving Models in Production

Warm-up and cold starts

A cold start is the extra time a fresh server instance needs before it answers at full speed, from starting the process through loading weights to its first few slow calls.

On this page 7
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A cold start is the extra time a brand-new server instance needs before it can answer requests at full speed.

The analogy you have already lived

A car engine on a freezing morning does not deliver full power the instant you turn the key. Oil has to warm and thin, moving parts need a moment to settle in. A car that has been running for twenty minutes responds instantly. A car started thirty seconds ago does not, even though both engines are identical.

A freshly started model server is the cold engine. It becomes the same server, though not immediately.

Why it exists

Model serving already established "load the model once, at startup". That advice is about correctness — never reload per request. This lesson is about what happens in those first moments after startup, before things settle down.

Starting a new server instance involves several slow steps that only happen once. This can be because of a deploy, a crash restart, or autoscaling adding capacity. The process starts. Code imports resolve. The model file loads from disk. Often, the first few real calls are genuinely slower too, before performance reaches its steady state.

How it works

new instance starts
      |
      v
  process starts, imports resolve        <- takes time
      |
      v
  model file read from disk              <- takes time, grows with model size
      |
      v
  first few real predictions             <- often slower than later ones
      |
      v
  STEADY STATE  <- this is what "fast" actually means

Every one of those steps already happened once before. That was back when you tested it on your laptop, with a server that had already been running for a while. That is exactly why cold starts are easy to miss, until real traffic meets a freshly started instance.

A real example you have seen

An app that feels sluggish for the very first search after you open it. Every search after that feels instant. The first call often pays a warm-up cost the rest do not.

Remember this

  • A cold start is the extra time a fresh instance needs before reaching its normal speed.
  • It includes starting the process, loading the model, and often genuinely slower first calls.
  • Testing on an already-running server hides this completely — you have to measure a truly fresh start to see it.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn joblib numpy

Measuring a real, small cold-start effect

warmup_timing.py
import time
import joblib
import numpy as np
from sklearn.linear_model import LogisticRegression

X = np.random.RandomState(0).uniform(0, 10, (300, 3))
y = (X.sum(axis=1) > 15).astype(int)
model = LogisticRegression().fit(X, y)
joblib.dump(model, "warm_model.joblib")

# Cold: the process has only recently started. This is the FIRST call after the
# model file is read from disk -- imports resolve, numpy's BLAS threads
# spin up, Python compiles a few code paths it has not run yet.
t0 = time.perf_counter()
loaded = joblib.load("warm_model.joblib")
load_time = time.perf_counter() - t0

row = np.array([[5.0, 5.0, 5.0]])

t0 = time.perf_counter()
loaded.predict_proba(row)
first_call = time.perf_counter() - t0

# Warm: the 50th call, everything above has already happened once.
times = []
for _ in range(50):
    t0 = time.perf_counter()
    loaded.predict_proba(row)
    times.append(time.perf_counter() - t0)

print(f"model file load             : {load_time*1000:.2f} ms")
print(f"first predict call (cold)   : {first_call*1000:.3f} ms")
print(f"predict call, calls 2-50 avg: {sum(times[1:])/len(times[1:])*1000:.3f} ms")
print(f"predict call, slowest of 50 : {max(times)*1000:.3f} ms")
print(f"predict call, fastest of 50 : {min(times)*1000:.3f} ms")
Output
model file load             : 6.58 ms
first predict call (cold)   : 0.197 ms
predict call, calls 2-50 avg: 0.041 ms
predict call, slowest of 50 : 0.160 ms
predict call, fastest of 50 : 0.036 ms

These are real wall-clock numbers from one run on one CPU — every number here will differ on your machine, and even between runs on the same machine. What is worth trusting is the shape: the first call ran roughly five times slower than the later average, for a tiny logistic regression doing almost no real work. That gap grows dramatically for a large neural network, especially on a GPU, for reasons this small CPU demo cannot reproduce — see the researcher section below.

Line-by-line walkthrough

joblib.load timing captures reading the file from disk and reconstructing the Python object — real cost that scales with model size, from milliseconds for this toy model to seconds for a large one.

The gap between "first predict call" and the "calls 2-50 average" is the actual warm-up effect: the first call pays a small one-time cost (memory allocation, code paths run for the first time) that later calls do not.

Common mistakes

Load-testing a server that has already been running for an hour, then trusting those numbers for a fresh deploy. This measures steady state, not what a real user hits seconds after a new instance starts serving traffic. Test a genuinely fresh instance too.

Scaling to zero for cost savings, without accounting for the wake-up cost. A CPU-only FastAPI server might wake up in under a second; a large model on a GPU can take tens of seconds to minutes, an eternity for the first user who hits it.

Treating the cold-start cost as fixed regardless of model size. A tiny logistic regression's cold start is barely measurable, as this demo shows. A multi-gigabyte transformer's is not — do not extrapolate one onto the other.

Try it yourself

Replace the logistic regression with a small RandomForestClassifier with 500 trees, and re-run. Compare the gap between the first call and the steady-state average against this lesson's logistic regression numbers.

What to learn next

Researcher — Mathematics and papers.

The full cold-start chain

For a containerised deployment, the complete cold path has more stages than the demo above can show on CPU alone:

  1. Scheduling — the orchestrator (Kubernetes, a serverless platform) finds a node with capacity.
  2. Image pull — the container image is downloaded, unless already cached on that node. For a multi-gigabyte image with CUDA libraries, this alone can take tens of seconds on a cold node.
  3. Process start — the runtime starts, imports resolve.
  4. Weight load — model weights are read from disk or object storage into memory (and onto the GPU, if applicable).
  5. Framework warm-up — the first forward pass, which for GPU frameworks often triggers CUDA kernel selection, cuDNN algorithm autotuning, and (if used) torch.compile or TensorRT graph compilation, all of which are one-time costs paid on the first real call.

Each stage dominates in different regimes: image pull dominates for large, uncached images; weight load dominates for very large models on fast storage; framework warm-up dominates for GPU-heavy models using compiled or autotuned execution paths.

Why GPU warm-up is much larger than CPU warm-up

cuDNN's autotuning mode benchmarks several convolution algorithms on the first call with a given input shape, selecting the fastest for subsequent calls with that same shape — a real, measured cost that can add hundreds of milliseconds to seconds on the first call, invisible on CPU because no equivalent autotuning step exists in typical CPU numerical libraries. torch.compile and TensorRT engine building are more extreme versions of the same trade: a large one-time compilation cost in exchange for lower steady-state latency, which is precisely why Triton's TensorRT backend pre-compiles ahead of deployment rather than on a live request.

Mitigations, in decreasing order of typical effect

  • A warm floor of instances that are never scaled to zero, absorbing traffic while cold instances start.
  • Pre-pulled images cached on nodes ahead of a scaling event, removing the image-pull stage from the critical path.
  • Loading weights from local fast storage (NVMe) rather than remote object storage, reducing the weight-load stage.
  • Synthetic warm-up requests sent to a new instance before it receives real traffic, deliberately paying the framework-warm-up cost against fake data rather than a real user's first request.
  • Scaling on queue depth or in-flight request count rather than CPU utilisation, since CPU is a lagging indicator for GPU-bound services and triggers scale-out too late for the cold-start time to matter — a point model serving's researcher section also makes.

Papers and systems

  • Chellappan et al., Serverless Cold Starts and Where To Find Them, EuroSys workshop papers document the general cold-start problem for the serverless case this lesson's GPU version is a heavier instance of.
  • NVIDIA's cuDNN developer guide documents autotuning behaviour and its first-call cost directly.
  • Google's Vertex AI and AWS SageMaker both publish documented guidance on minimum warm-instance counts specifically to bound cold-start-driven tail latency for GPU-backed endpoints.

What to learn next