Scaling and Traffic Management

Autoscaling on the right metric

Autoscaling adds and removes machines automatically, but only helps if it watches the number that actually predicts trouble.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. The trap almost everyone falls into first
  5. How it works
  6. Where you have already benefited from this
  7. The honest part
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Autoscaling means a program adds or removes machines by itself, based on one number it watches.

The analogy you have already lived

You have stood in line at a chai stall on a busy evening. The stall owner does not check how hot the stove is before calling in a second pair of hands. The owner looks at the line — how many people are waiting, and how long they have been standing there.

That line length is the only number that tells the owner whether help is needed right now. The stove could be roaring and the line could still be short. The stove could be gentle and the line could still be out the door.

Why it exists

Before autoscaling, someone had to guess how many machines a service would need, and buy that many in advance.

Guess too low, and the service falls over during a sale, a cricket final, or exam results day. Guess too high, and you pay for idle machines every single night.

Autoscaling replaces the guess with a rule. Watch one number, and add or remove machines when it crosses a line you chose. The whole system now depends on that one choice — which number to watch.

The trap almost everyone falls into first

The obvious-looking choice is CPU usage — how busy the processor is, as a percentage.

For an inference server, this is usually the wrong signal. Most of a request's time is not spent computing. It is spent waiting — for a GPU to finish, for a network call to a database. Sometimes it is waiting for another service to answer.

Waiting looks idle to the operating system. A server can have forty requests queued up, genuinely struggling, while its CPU usage sits at eight percent.

How it works

   requests arrive
        |
        v
   [ waiting line ]   <- length checked every few seconds
        |
        v
   [ 4 workers, all busy waiting on a GPU call ]

   CPU usage:     8%      <- looks calm
   line length:  47       <- the real story

The fix is to watch queue depth — how many requests are waiting their turn. Or requests-in-flight — how many are currently being handled. Both describe the line, not the stove.

Where you have already benefited from this

A food delivery app during dinner rush adds more order-matching servers as orders pile up. It does not wait for its processors to get warm. A bank's fraud-check service scales up around a big sale, watching how many payments are waiting to be scored. That is why your UPI payment does not hang for ten seconds while everyone else checks out too.

The honest part

Picking the right metric is a decision, not a default setting. Most autoscaling tools ship with CPU-based scaling turned on because it is easy to measure everywhere. Easy to measure is not the same as useful to measure.

Remember this

  • Autoscaling reacts to one chosen number — pick a number that tracks real waiting, not raw busyness.
  • CPU usage can look calm while a service is genuinely overloaded. This happens most often when most of the time is spent waiting on something else.
  • Queue depth or requests-in-flight describes the line of waiting work directly, which is what autoscaling is meant to relieve.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install psutil

A queue that fills up while the CPU stays quiet

This simulates a small inference service: four worker threads, each handling one request at a time. Each request takes 50ms because it is "waiting on a GPU" — a time.sleep call, which uses no CPU at all. A burst of 60 requests lands almost at once.

metric_demo.py
import queue
import threading
import time

import psutil

NUM_WORKERS = 4
SERVICE_TIME = 0.05  # seconds per request; stands in for a GPU/network call

work = queue.Queue()
depth_samples = []
stop = False


def worker():
    while True:
        item = work.get()
        if item is None:
            work.task_done()
            return
        time.sleep(SERVICE_TIME)  # waiting on the "model", not using the CPU
        work.task_done()


def sampler():
    proc = psutil.Process()
    proc.cpu_percent()  # first call primes the measurement
    while not stop:
        time.sleep(0.02)
        depth_samples.append((work.qsize(), proc.cpu_percent()))


threads = [threading.Thread(target=worker) for _ in range(NUM_WORKERS)]
for t in threads:
    t.start()
sample_thread = threading.Thread(target=sampler)
sample_thread.start()

for _ in range(60):        # a burst: 60 requests land in well under a second
    work.put(object())

work.join()
stop = True
sample_thread.join()
for _ in range(NUM_WORKERS):
    work.put(None)
for t in threads:
    t.join()

peak_backlog = max(d for d, _ in depth_samples)
peak_cpu = max(c for _, c in depth_samples)
print(f"samples taken:      {len(depth_samples)}")
print(f"peak queue backlog: {peak_backlog} requests waiting")
print(f"peak process CPU%:  {peak_cpu:.1f}")
Output
samples taken:      37
peak queue backlog: 56 requests waiting
peak process CPU%:  0.0

This machine's CPU% is exactly 0.0 because the simulated work is pure waiting. A time.sleep call uses no processor time at all. Your own run will match the backlog number closely, since it does not depend on hardware speed. CPU% may show a small nonzero value on a busier machine, but it will stay far below what a 56-deep backlog deserves.

Walking through it

time.sleep(SERVICE_TIME) inside the worker stands in for the part of a real request that is out of the CPU's hands. Think of a GPU call, or a network round trip. This is deliberate: it makes the CPU-vs-queue gap visible without needing real hardware.

work.qsize() is the number of requests waiting, not the number being worked on. That waiting count is what an autoscaler should watch. It answers "is work piling up faster than we can clear it?" directly.

proc.cpu_percent() reports how busy this process's threads look to the operating system. The first call always returns 0.0 by design — psutil's own documentation notes this. That is why it is called once before sampling begins.

Common mistakes

Scaling on CPU for a service that mostly waits. Any service spending its time on GPU calls, database lookups, or downstream APIs will under-report load on CPU. Check what the service actually spends its time doing before choosing a metric for it.

Averaging over too long a window. A five-minute average of queue depth smooths out a two-minute burst completely. By the time the average catches up, the burst is over and users have already seen it.

Watching a metric that stays flat regardless of load. Memory usage on a service that loads one fixed model at startup barely moves. It stays about the same whether serving one request a second or a thousand. A flat metric cannot signal anything.

Scaling on a total instead of a per-instance number. "200 requests waiting" means something different across 2 machines than across 20. Divide by instance count, or scale on a per-instance figure, so the target line stays meaningful as the fleet grows.

Try it yourself

Change SERVICE_TIME to 0.005 (10x faster) and rerun. Watch the peak backlog shrink, because the workers now clear the same burst before it can pile up. This is the mechanic an autoscaler is reacting to. Backlog is a race between arrival speed and service speed, not a fixed property of the traffic.

What to learn next

Researcher — Mathematics and papers.

Autoscaling as feedback control

An autoscaler is a control loop: observe a metric, compare it to a target, adjust capacity, wait, repeat. Hellerstein, Diao, Parekh and Tilbury's Feedback Control of Computing Systems (2004) frames this formally, and the classic failure modes of control theory apply directly — oscillation from reacting too aggressively, and steady-state error from reacting too weakly.

The relationship between arrival rate, wait time and the number of items in a system is Little's Law, covered with its full derivation in the researcher section of model serving. Autoscaling on queue depth is, in that framing, autoscaling on $L$ directly — the quantity the law says predicts wait time, rather than on a proxy for it.

Concurrency as the target metric

Knative's autoscaler (KPA) targets concurrency — the number of requests each instance is handling at once — rather than CPU. An operator sets a target concurrency per instance; the autoscaler solves for the instance count that keeps observed concurrency near that target. This is architecturally the same idea as the queue-depth signal above, generalised to a per-instance rate rather than a raw count.

Kubernetes' Horizontal Pod Autoscaler supports custom and external metrics for exactly this reason — CPU and memory are the defaults, not the recommendation, for request-serving workloads.

Reactive versus predictive scaling

Everything above is reactive: it responds after the metric moves, and pays the lag described in autoscaling lag and headroom. Predictive autoscaling instead forecasts tomorrow's traffic from historical patterns and pre-provisions ahead of it. Netflix's Scryer engine, documented in Netflix's own 2013 technology blog post "Scryer: Netflix's Predictive Auto Scaling Engine," combined pattern recognition with reactive scaling as a safety net. The combination exists specifically to absorb the minutes-long lag of booting a new instance ahead of known daily and weekly traffic cycles.

Survey literature

  • Lorido-Botran, Miguel-Alonso and Lozano, A Review of Auto-scaling Techniques for Elastic Applications in Cloud Environments, Journal of Grid Computing, 2014 — a broad taxonomy of reactive, predictive and hybrid approaches, and the metrics each family relies on.
  • Al-Dhuraibi, Paraiso, Djarallah and Merle, Elasticity in Cloud Computing: State of the Art and Research Challenges, IEEE Transactions on Cloud Computing, 2018 — surveys the gap between coarse infrastructure metrics and application-level signals like queue depth and concurrency.
  • Hellerstein, Diao, Parekh and Tilbury, Feedback Control of Computing Systems, Wiley, 2004 — the control-theory foundation underneath every autoscaler's threshold and cooldown logic.

What to learn next