GPUs: Memory, Scheduling and Cost

Serverless GPU platforms

Serverless GPU platforms give you a GPU only for the seconds you actually use one, and hand it back afterwards — in exchange for a real, measurable cold-start cost.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A serverless GPU platform gives you a GPU only for the moments you actually use one, and takes it back afterwards. You never own or rent a machine full-time.

The analogy you have already lived

You do not own a car to get from home to a friend's house once a month. You call a taxi, it arrives, takes you there, and leaves. You pay for the ride you actually took. Not for a car sitting in your driveway the other twenty-nine days of the month.

The trade-off is real: a taxi takes a few minutes to arrive. A car in your driveway is ready the instant you want it. Serverless GPUs make exactly this trade, for computing instead of transport.

Why it exists

Renting a GPU by the hour, all day, is wasteful if your actual usage is bursty. Busy for a few minutes, then quiet for a long stretch. Owning or renting a full-time GPU for that pattern means paying for a lot of idle time.

A serverless GPU platform starts a GPU machine for you the moment a request arrives. It runs your model, then shuts back down afterwards. You are billed for the seconds actually used, not for the hours in between.

How it works

   no requests: NOTHING is running, NOTHING is being billed
              |
   a request finally arrives
              |
              v
   a GPU machine is started, JUST for this
   (this step takes real time -- see below)
              |
              v
   your model loads, answers the request
              |
              v
   quiet again: the machine is shut down, billing stops

The step marked "takes real time" is the entire trade-off this lesson is about. It has a name: a cold start.

A real example you have seen

Ride-hailing apps show you an estimated wait before your driver arrives. A real, honest number — the car was not sitting outside your door already. Serverless GPU platforms have an equivalent number for the very first request after a quiet period. It is worth knowing before you build around one.

The honest part

The cold-start delay is real, and it can be large. Seconds, sometimes longer, depending on how big the model is. For a chatbot where a user is watching the screen, waiting several seconds for the first reply of the day is a genuinely bad experience. For a background job with nobody watching, the same delay is often completely irrelevant. Which one your situation is decides whether serverless is the right fit.

Remember this

  • Serverless GPUs bill you only for actual use, not for idle time in between.
  • The trade-off is a real, measurable cold-start delay on the first request after a quiet period.
  • Whether that trade-off is worth it depends entirely on whether someone is waiting for the answer.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch transformers

Measuring a real cold start against a real warm one

cold_vs_warm.py
import subprocess
import sys
import time

COLD_SCRIPT = '''
import time
start = time.perf_counter()
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("distilgpt2")
tok.pad_token = tok.eos_token
model = AutoModelForCausalLM.from_pretrained("distilgpt2").to("cuda")
enc = tok("Hello", return_tensors="pt").to("cuda")
with torch.no_grad():
    model.generate(**enc, max_new_tokens=5, do_sample=False, pad_token_id=tok.eos_token_id)
print(time.perf_counter() - start)
'''
with open("cold_worker.py", "w") as f:
    f.write(COLD_SCRIPT)

# A brand-new process: imports, loads the model, answers ONE request.
# This stands in for a serverless platform's cold start -- a fresh
# environment with nothing already loaded.
start = time.perf_counter()
result = subprocess.run([sys.executable, "cold_worker.py"], capture_output=True, text=True)
wall_time = time.perf_counter() - start
reported_time = float(result.stdout.strip().splitlines()[-1])
print(f"cold start (new process, import + load model + first answer): {reported_time:.2f}s")
print(f"  (measured from outside the process: {wall_time:.2f}s, includes Python startup itself)")

# A WARM request: the model is already loaded in THIS process.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("distilgpt2")
tok.pad_token = tok.eos_token
model = AutoModelForCausalLM.from_pretrained("distilgpt2").to("cuda")
enc = tok("Hello", return_tensors="pt").to("cuda")
with torch.no_grad():
    model.generate(**enc, max_new_tokens=5, do_sample=False, pad_token_id=tok.eos_token_id)  # warm-up

start = time.perf_counter()
with torch.no_grad():
    model.generate(**enc, max_new_tokens=5, do_sample=False, pad_token_id=tok.eos_token_id)
warm_time = time.perf_counter() - start
print(f"warm request (model already loaded, answering again): {warm_time*1000:.1f} ms")
print(f"cold start is roughly {reported_time/warm_time:.0f}x slower than a warm request")
Output
cold start (new process, import + load model + first answer): 8.63s
  (measured from outside the process: 9.38s, includes Python startup itself)
warm request (model already loaded, answering again): 32.9 ms
cold start is roughly 262x slower than a warm request

This is a genuine, measured cold start. A brand-new Python process imports libraries and loads a small model from scratch, measured against a warm request in a process that already had everything ready. A real serverless platform's cold start includes more steps this simulation cannot show on one machine: starting a new virtual machine or container, pulling a container image, attaching a GPU. Real cold starts are frequently slower than this number, not faster — especially for models much larger than the small one used here.

Line-by-line walkthrough

A fresh subprocess for the "cold" measurement. This is the honest part of simulating a cold start on one machine. A brand-new Python process pays the full cost of importing libraries and loading the model from disk, exactly like a fresh serverless container would — minus the container and networking layers a real platform adds on top.

The warm measurement reuses an already-loaded model. This mirrors what a serverless platform's "keep warm" feature does. The second and later requests to an already-running instance skip the loading step entirely — the entire reason the warm number is so much smaller.

The roughly 262x gap. This specific number is a real measurement, from a small model, on this exact machine. Treat the scale of the gap, not the specific multiplier, as the transferable lesson. A larger, real production model widens this gap further. Loading time grows with model size, while a warm forward pass does not grow nearly as fast.

Common mistakes

Assuming serverless GPU cost is only about the compute. Cold starts are wasted, billed time. You often pay for the seconds spent loading, not only the seconds spent actually answering. A platform with frequent cold starts can cost more than the sticker price per second suggests.

Ignoring "keep warm" settings that cost real money. Most serverless GPU platforms let you pay to keep an instance warm between requests. That trades the cold-start delay away for a cost closer to always-on. This defeats a meaningful part of the cost benefit if set too aggressively for genuinely low, bursty traffic.

Choosing serverless for a use case with a hard latency requirement. Say a user is actively waiting on the very first request of the day, and that request can take several seconds. Serverless may be the wrong shape entirely — see self-hosting vs API: the break-even point for the always-on alternative.

Not testing cold-start time with your actual model. The gap scales with model size and the specific platform's container-and-VM startup overhead — both vary enormously. Measure it directly with your real model before committing to an architecture around it.

Try it yourself

Change distilgpt2 to a larger model your machine can handle, and rerun. Watch how much the cold-start number grows compared to the warm one. Consider what that gap means for a model many times larger still.

What to learn next

Researcher — Mathematics and papers.

Decomposing a real cold start

A production serverless GPU cold start is the sum of several stages this lesson's local simulation cannot fully reproduce on one machine:

$$T_{\text{cold}} = T_{\text{provision}} + T_{\text{image pull}} + T_{\text{container start}} + T_{\text{driver init}} + T_{\text{model load}} + T_{\text{first inference}}$$

  • $T_{\text{provision}}$ — acquiring a physical or virtual machine with an attached GPU, which can itself involve waiting for spare capacity.
  • $T_{\text{image pull}}$ — downloading a container image, often the largest term for large ML images bundling CUDA, frameworks and model weights, unless the platform caches images on the node already.
  • $T_{\text{model load}}$ — reading weights from storage into GPU memory; for large models pulled from remote object storage rather than local disk, this alone can dominate the entire cold start.

Platforms differentiate primarily by attacking these terms: pre-warmed pools of idle GPU instances remove $T_{\text{provision}}$ almost entirely at a cost of paying for genuine idle capacity; snapshotting a fully-initialised process (CRIU-based checkpoint-restore, used by several serverless ML platforms) can skip past $T_{\text{driver init}}$ and much of $T_{\text{model load}}$ by resuming a previously-initialised memory image rather than reconstructing it from scratch; and local NVMe caching of model weights on the node reduces $T_{\text{model load}}$ compared to pulling from remote storage on every cold start.

The scale-to-zero trade-off, formally

Scale-to-zero (shutting down completely between requests) minimises idle cost but maximises cold-start frequency. A "keep warm for $w$ seconds after the last request" policy trades a bounded extra idle cost — at most $w$ seconds of unused billing per idle gap — against a reduced cold-start rate, whenever request inter-arrival times are shorter than $w$ often enough to matter. The optimal $w$ for a given traffic pattern is a real queueing-theory tuning problem, not a universal constant, and depends on the ratio between cold-start cost and idle cost specifically for that workload.

Where serverless GPU platforms differ from serverless CPU functions

Traditional serverless (AWS Lambda, Cloud Functions) popularised the model for stateless, lightweight, CPU-bound functions with startup times often well under a second. GPU serverless inherits the same conceptual model, but faces structurally larger cold starts. GPU drivers, CUDA contexts and multi-gigabyte model weights are heavier state to initialise than a typical CPU function's dependencies. Intuitions and tuning guidance carried over directly from CPU serverless experience routinely underestimate real GPU cold-start cost.

Reading

  • Google Cloud, Cloud Run GPU cold start documentation and benchmarks — a concrete, current reference point for real-world cold-start figures
  • Modal, Runpod, Replicate — commercial serverless GPU platforms; their own published cold-start benchmarks are a useful cross-check against any single measurement, since they vary by model size, region and image caching strategy

What to learn next