Scaling and Traffic Management
Scale to zero, and what it costs you
Scaling to zero saves money while nobody is asking, and makes the first person to ask wait for a cold start.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Scale to zero means shutting a service down completely when nobody is using it. It starts fresh the moment someone does.
The analogy you have already lived
You have turned off the gas stove between customers at a quiet roadside dhaba. It saves gas. It also means the next customer waits a little longer. The pan has to heat up again before anything can cook.
Keeping the stove lit all day costs gas every hour, busy or not. Turning it off between customers costs time instead, every time someone new arrives.
Why it exists
Every machine that is running costs money for every hour it runs, whether or not anyone uses it. A service that gets one request a day, but runs 24 hours anyway, is mostly paying for doing nothing.
Scale to zero removes that waste. When no requests have arrived for a while, the service shuts down entirely — no machine, no bill. When a request finally arrives, a machine is started, the model is loaded, and only then is the request answered.
The price: a cold start
That first request after silence has to wait for the whole service to boot. A machine is assigned, an image downloaded, a process started, a model file read into memory. This wait is called a cold start. It can run from a few hundred milliseconds to tens of seconds. The size of the model and the container image both matter.
How it works
quiet for a while
|
v
[ service shut down completely, costing nothing ]
|
a request arrives
|
v
[ boot machine -> load model -> answer ] <- the cold start, felt by this one user
|
v
[ service stays warm for the next few minutes, answering instantly ]
|
v
quiet again -> shuts down -> repeatA real example you have seen
A rarely used internal tool at a small company is a natural fit. Think of a document search bot the team checks a few times a day. Nobody minds a two-second wait for something used twice a day. A payment-confirmation service that every customer touches on every order is the opposite case. Even a half-second cold start, multiplied across thousands of orders, is a real cost in frustrated customers.
The honest part
Scale to zero is a trade, not a free win. It swaps a certain cost — the always-on bill — for an uncertain one. That uncertain cost is how often a real user gets stuck waiting, and how much that is worth to you. The right answer depends entirely on how your specific traffic is shaped, not on a general rule.
Remember this
- Scale to zero costs nothing while idle. It costs a cold start — real wait time — the moment traffic returns.
- The bigger the model and the container image, the longer that cold start tends to be.
- It fits bursty, occasional, low-stakes traffic far better than it fits anything a user is actively waiting on.
What to learn next
- Load balancing inference traffic — once more than one instance exists, warm or cold, how requests are routed to it.
- Autoscaling lag and headroom — the general version of the cold-start problem, for scaling up rather than scaling from zero.
- Model serving — the service being scaled to zero here, and the health check that tells you when it has finished booting.
Developer — Code and libraries.
Setup
pip install scikit-learn pandas joblib psutilMeasuring a real cold start versus a real warm reload
This trains a small model, then measures the honest difference. Starting a brand-new Python process and loading everything from scratch, against reusing a process that is already warmed up.
import joblib
import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
rng = np.random.RandomState(0)
n = 400
X = pd.DataFrame({
"income": rng.uniform(5, 80, n).round(1),
"years": rng.uniform(0, 10, n).round(1),
"age": rng.randint(21, 65, n).astype(float),
})
score = 0.05 * X["income"] + 0.35 * X["years"] + 0.01 * X["age"] - 3.5
y = (score + rng.normal(0, 0.8, n) > 0).astype(int)
model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000)).fit(X, y)
joblib.dump(model, "model.joblib")import subprocess
import sys
import time
import joblib
import numpy as np
# Cold: a fresh Python process, importing every library from scratch.
code = "import joblib, pandas as pd; joblib.load('model.joblib')"
t = time.perf_counter()
subprocess.run([sys.executable, "-c", code], check=True)
cold_ms = (time.perf_counter() - t) * 1000
# Warm: the interpreter and libraries are already loaded, only the file read happens.
warm_times = []
for _ in range(5):
t = time.perf_counter()
joblib.load("model.joblib")
warm_times.append((time.perf_counter() - t) * 1000)
warm_ms = sum(warm_times) / len(warm_times)
print(f"cold start (new process): {cold_ms:7.1f} ms")
print(f"warm reload (file only): {warm_ms:7.2f} ms")
print(f"cold is roughly {cold_ms / warm_ms:.0f}x slower")
# --- the cost side: a simulation with a fixed seed, not a live measurement ---
PRICE_PER_HOUR = 0.05 # a made-up small-instance price, for illustration only
REQUESTS_PER_DAY = 300
IDLE_TIMEOUT_MIN = 10 # scale-to-zero shuts down after this much silence
rng = np.random.RandomState(3)
arrivals = np.sort(rng.uniform(0, 24 * 60, REQUESTS_PER_DAY)) # minutes since midnight
# Merge each [arrival, arrival + timeout] window with the next if they overlap.
# A merged run of windows = one "instance stayed up" stretch = one cold start.
cold_starts = 1
active_minutes = 0.0
window_start, window_end = arrivals[0], arrivals[0] + IDLE_TIMEOUT_MIN
for a in arrivals[1:]:
if a <= window_end:
window_end = a + IDLE_TIMEOUT_MIN
else:
active_minutes += window_end - window_start
window_start, window_end = a, a + IDLE_TIMEOUT_MIN
cold_starts += 1
active_minutes += window_end - window_start
always_on_cost = PRICE_PER_HOUR * 24
scale_to_zero_cost = PRICE_PER_HOUR * (active_minutes / 60)
print(f"\nsimulated cold starts in one day: {cold_starts}")
print(f"instance-minutes billed: {active_minutes:.0f} of 1440")
print(f"always-on cost: ${always_on_cost:.2f} / day")
print(f"scale-to-zero cost: ${scale_to_zero_cost:.2f} / day")cold start (new process): 1099.6 ms warm reload (file only): 149.28 ms cold is roughly 7x slower simulated cold starts in one day: 35 instance-minutes billed: 1235 of 1440 always-on cost: $1.20 / day scale-to-zero cost: $1.03 / day
The millisecond numbers in the first block are real measurements from this machine, taken right now. Yours will differ, with your own CPU, disk, and running programs. The gap between cold and warm — several times slower — is the part that reliably reproduces.
The second block is a simulation, not a measurement. 300 request times were drawn from a fixed random seed, and the arithmetic that follows is exact given those numbers. It illustrates the mechanic, not a real bill.
Walking through it
The subprocess trick. Running joblib.load inside a brand-new python -c "..." process is the honest way to measure a cold start on one machine. It pays the real cost of starting an interpreter and importing every library again. That is exactly what a scale-to-zero platform does when it boots a fresh instance.
Merging overlapping windows. A naive estimate — multiply request count by idle timeout — badly overcounts. A busy period is one continuous "stayed warm" stretch, not several separate ones. The merge loop is what turns 300 individual arrival times into a small number of real active-and-idle windows.
35 cold starts, not 300. Even with a fairly short 10-minute idle timeout, most requests arrived close enough together to reuse an already-warm instance. Cold starts happen at the start of each isolated burst of traffic, not on every request.
Common mistakes
Setting the idle timeout too short. A very short timeout — say, 30 seconds — turns nearly every request into its own cold start. The gap between typical requests outlasts it entirely. Measure your actual inter-request gaps before choosing a number.
Ignoring cold starts because "the average is fine." Averages hide the worst case. The users who matter here are the ones who arrive right after a quiet stretch. It is a small fraction of requests, but a real and repeatable one.
Assuming savings scale to any traffic level. At high enough request volume, windows merge so often that the service is almost always warm. Scale-to-zero then saves little, while still risking occasional cold starts. It earns its keep at low, bursty volume — check which one you have.
Try it yourself
Change IDLE_TIMEOUT_MIN to 2 and rerun. Watch cold_starts rise sharply as more gaps between requests exceed the shorter timeout. The same 300 requests now trigger far more boots.
What to learn next
- Load balancing inference traffic — once more than one instance exists, warm or cold, how requests are routed to it.
- Autoscaling lag and headroom — the general version of the cold-start problem, for scaling up rather than scaling from zero.
- Model serving — the service being scaled to zero here, and the health check that tells you when it has finished booting.
Researcher — Mathematics and papers.
The cost side as an economic trade-off
Let $p$ be the hourly price of one running instance, $\lambda$ the average request rate, and $\tau$ the idle timeout. In a Poisson arrival process with rate $\lambda$, the expected number of separate active windows per unit time — each one a cold start — falls as $\lambda$ or $\tau$ rises. More requests then arrive within any given $\tau$-length gap, and get merged into the same window. This is why the simulation above shows diminishing cold-start counts as traffic gets busier: it is a direct consequence of the arrival process, not an artefact of the specific numbers chosen.
The always-on cost is fixed at $24p$ per day regardless of traffic. The scale-to-zero cost approaches it from below as $\lambda$ grows, and the crossover point — where scale-to-zero stops being meaningfully cheaper — depends on $\tau$, $\lambda$, and how expensive a cold start is to the business, not on price alone.
Mitigating cold starts without giving up the savings
- Provisioned concurrency (AWS Lambda) or a minimum instance count (Google Cloud Run, Knative) keeps a small, non-zero floor warm at all times — converting scale-to-zero into "scale to a small headroom," trading some savings back for a latency guarantee.
- Snapshotting a warm process and restoring it, rather than booting from nothing, is the mechanism behind fast serverless cold starts on platforms like Firecracker microVMs — the OS and interpreter state is restored rather than reconstructed.
- Splitting the model load from the process start bounds the user-visible wait to only the model-load portion of the cold start, not the whole boot chain. This means keeping a small always-on router that forwards to a scale-to-zero backend only after the backend signals readiness.
Where the research sits
Wang et al., Peeking Behind the Curtains of Serverless Platforms, USENIX ATC 2018, measured cold-start behaviour across AWS Lambda, Azure Functions and Google Cloud Functions empirically. They found cold-start latency dominated by language runtime and package size — directly consistent with the subprocess-versus-warm-reload gap measured above, at a much larger scale.
What to learn next
- Load balancing inference traffic — once more than one instance exists, warm or cold, how requests are routed to it.
- Autoscaling lag and headroom — the general version of the cold-start problem, for scaling up rather than scaling from zero.
- Model serving — the service being scaled to zero here, and the health check that tells you when it has finished booting.