Scaling and Traffic Management
Autoscaling on the right metric
Autoscaling adds and removes machines automatically, but only helps if it watches the number that actually predicts trouble.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Autoscaling means a program adds or removes machines by itself, based on one number it watches.
The analogy you have already lived
You have stood in line at a chai stall on a busy evening. The stall owner does not check how hot the stove is before calling in a second pair of hands. The owner looks at the line — how many people are waiting, and how long they have been standing there.
That line length is the only number that tells the owner whether help is needed right now. The stove could be roaring and the line could still be short. The stove could be gentle and the line could still be out the door.
Why it exists
Before autoscaling, someone had to guess how many machines a service would need, and buy that many in advance.
Guess too low, and the service falls over during a sale, a cricket final, or exam results day. Guess too high, and you pay for idle machines every single night.
Autoscaling replaces the guess with a rule. Watch one number, and add or remove machines when it crosses a line you chose. The whole system now depends on that one choice — which number to watch.
The trap almost everyone falls into first
The obvious-looking choice is CPU usage — how busy the processor is, as a percentage.
For an inference server, this is usually the wrong signal. Most of a request's time is not spent computing. It is spent waiting — for a GPU to finish, for a network call to a database. Sometimes it is waiting for another service to answer.
Waiting looks idle to the operating system. A server can have forty requests queued up, genuinely struggling, while its CPU usage sits at eight percent.
How it works
requests arrive
|
v
[ waiting line ] <- length checked every few seconds
|
v
[ 4 workers, all busy waiting on a GPU call ]
CPU usage: 8% <- looks calm
line length: 47 <- the real storyThe fix is to watch queue depth — how many requests are waiting their turn. Or requests-in-flight — how many are currently being handled. Both describe the line, not the stove.
Where you have already benefited from this
A food delivery app during dinner rush adds more order-matching servers as orders pile up. It does not wait for its processors to get warm. A bank's fraud-check service scales up around a big sale, watching how many payments are waiting to be scored. That is why your UPI payment does not hang for ten seconds while everyone else checks out too.
The honest part
Picking the right metric is a decision, not a default setting. Most autoscaling tools ship with CPU-based scaling turned on because it is easy to measure everywhere. Easy to measure is not the same as useful to measure.
Remember this
- Autoscaling reacts to one chosen number — pick a number that tracks real waiting, not raw busyness.
- CPU usage can look calm while a service is genuinely overloaded. This happens most often when most of the time is spent waiting on something else.
- Queue depth or requests-in-flight describes the line of waiting work directly, which is what autoscaling is meant to relieve.
What to learn next
- Autoscaling lag and headroom — why picking the right metric is not enough by itself.
- Model serving — the service this chapter assumes you already know how to build.
- Load balancing inference traffic — once you have several machines, how requests find the right one.
Developer — Code and libraries.
Setup
pip install psutilA queue that fills up while the CPU stays quiet
This simulates a small inference service: four worker threads, each handling one request at a time. Each request takes 50ms because it is "waiting on a GPU" — a time.sleep call, which uses no CPU at all. A burst of 60 requests lands almost at once.
import queue
import threading
import time
import psutil
NUM_WORKERS = 4
SERVICE_TIME = 0.05 # seconds per request; stands in for a GPU/network call
work = queue.Queue()
depth_samples = []
stop = False
def worker():
while True:
item = work.get()
if item is None:
work.task_done()
return
time.sleep(SERVICE_TIME) # waiting on the "model", not using the CPU
work.task_done()
def sampler():
proc = psutil.Process()
proc.cpu_percent() # first call primes the measurement
while not stop:
time.sleep(0.02)
depth_samples.append((work.qsize(), proc.cpu_percent()))
threads = [threading.Thread(target=worker) for _ in range(NUM_WORKERS)]
for t in threads:
t.start()
sample_thread = threading.Thread(target=sampler)
sample_thread.start()
for _ in range(60): # a burst: 60 requests land in well under a second
work.put(object())
work.join()
stop = True
sample_thread.join()
for _ in range(NUM_WORKERS):
work.put(None)
for t in threads:
t.join()
peak_backlog = max(d for d, _ in depth_samples)
peak_cpu = max(c for _, c in depth_samples)
print(f"samples taken: {len(depth_samples)}")
print(f"peak queue backlog: {peak_backlog} requests waiting")
print(f"peak process CPU%: {peak_cpu:.1f}")samples taken: 37 peak queue backlog: 56 requests waiting peak process CPU%: 0.0
This machine's CPU% is exactly 0.0 because the simulated work is pure waiting. A time.sleep call uses no processor time at all. Your own run will match the backlog number closely, since it does not depend on hardware speed. CPU% may show a small nonzero value on a busier machine, but it will stay far below what a 56-deep backlog deserves.
Walking through it
time.sleep(SERVICE_TIME) inside the worker stands in for the part of a real request that is out of the CPU's hands. Think of a GPU call, or a network round trip. This is deliberate: it makes the CPU-vs-queue gap visible without needing real hardware.
work.qsize() is the number of requests waiting, not the number being worked on. That waiting count is what an autoscaler should watch. It answers "is work piling up faster than we can clear it?" directly.
proc.cpu_percent() reports how busy this process's threads look to the operating system. The first call always returns 0.0 by design — psutil's own documentation notes this. That is why it is called once before sampling begins.
Common mistakes
Scaling on CPU for a service that mostly waits. Any service spending its time on GPU calls, database lookups, or downstream APIs will under-report load on CPU. Check what the service actually spends its time doing before choosing a metric for it.
Averaging over too long a window. A five-minute average of queue depth smooths out a two-minute burst completely. By the time the average catches up, the burst is over and users have already seen it.
Watching a metric that stays flat regardless of load. Memory usage on a service that loads one fixed model at startup barely moves. It stays about the same whether serving one request a second or a thousand. A flat metric cannot signal anything.
Scaling on a total instead of a per-instance number. "200 requests waiting" means something different across 2 machines than across 20. Divide by instance count, or scale on a per-instance figure, so the target line stays meaningful as the fleet grows.
Try it yourself
Change SERVICE_TIME to 0.005 (10x faster) and rerun. Watch the peak backlog shrink, because the workers now clear the same burst before it can pile up. This is the mechanic an autoscaler is reacting to. Backlog is a race between arrival speed and service speed, not a fixed property of the traffic.
What to learn next
- Autoscaling lag and headroom — why picking the right metric is not enough by itself.
- Model serving — the service this chapter assumes you already know how to build.
- Load balancing inference traffic — once you have several machines, how requests find the right one.
Researcher — Mathematics and papers.
Autoscaling as feedback control
An autoscaler is a control loop: observe a metric, compare it to a target, adjust capacity, wait, repeat. Hellerstein, Diao, Parekh and Tilbury's Feedback Control of Computing Systems (2004) frames this formally, and the classic failure modes of control theory apply directly — oscillation from reacting too aggressively, and steady-state error from reacting too weakly.
The relationship between arrival rate, wait time and the number of items in a system is Little's Law, covered with its full derivation in the researcher section of model serving. Autoscaling on queue depth is, in that framing, autoscaling on $L$ directly — the quantity the law says predicts wait time, rather than on a proxy for it.
Concurrency as the target metric
Knative's autoscaler (KPA) targets concurrency — the number of requests each instance is handling at once — rather than CPU. An operator sets a target concurrency per instance; the autoscaler solves for the instance count that keeps observed concurrency near that target. This is architecturally the same idea as the queue-depth signal above, generalised to a per-instance rate rather than a raw count.
Kubernetes' Horizontal Pod Autoscaler supports custom and external metrics for exactly this reason — CPU and memory are the defaults, not the recommendation, for request-serving workloads.
Reactive versus predictive scaling
Everything above is reactive: it responds after the metric moves, and pays the lag described in autoscaling lag and headroom. Predictive autoscaling instead forecasts tomorrow's traffic from historical patterns and pre-provisions ahead of it. Netflix's Scryer engine, documented in Netflix's own 2013 technology blog post "Scryer: Netflix's Predictive Auto Scaling Engine," combined pattern recognition with reactive scaling as a safety net. The combination exists specifically to absorb the minutes-long lag of booting a new instance ahead of known daily and weekly traffic cycles.
Survey literature
- Lorido-Botran, Miguel-Alonso and Lozano, A Review of Auto-scaling Techniques for Elastic Applications in Cloud Environments, Journal of Grid Computing, 2014 — a broad taxonomy of reactive, predictive and hybrid approaches, and the metrics each family relies on.
- Al-Dhuraibi, Paraiso, Djarallah and Merle, Elasticity in Cloud Computing: State of the Art and Research Challenges, IEEE Transactions on Cloud Computing, 2018 — surveys the gap between coarse infrastructure metrics and application-level signals like queue depth and concurrency.
- Hellerstein, Diao, Parekh and Tilbury, Feedback Control of Computing Systems, Wiley, 2004 — the control-theory foundation underneath every autoscaler's threshold and cooldown logic.
What to learn next
- Autoscaling lag and headroom — why picking the right metric is not enough by itself.
- Model serving — the service this chapter assumes you already know how to build.
- Load balancing inference traffic — once you have several machines, how requests find the right one.