GPUs: Memory, Scheduling and Cost
Running on spot and preemptible GPUs
Spot GPUs cost much less than normal, in exchange for the provider being allowed to take the machine back with little warning. Checkpointing is what makes that survivable.
- 9 min read
- 3 reading levels
- Published
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A spot GPU costs far less than a normal one. In exchange, the provider can take it back at short notice.
The analogy you have already lived
A waitlisted train ticket costs less than a confirmed one. You get a seat, right up until a confirmed passenger needs it. Then you are asked to give it up, sometimes with very little warning.
Say your journey does not depend on that exact seat on that exact train — you could catch a later one. A waitlisted ticket is then a genuinely good deal. If you absolutely have to arrive at a specific time, it is a real risk. One you are choosing to take on. A spot GPU is that same trade, made with a cloud provider instead of a railway.
Why it exists
Cloud providers have spare GPU capacity sitting unused at any given moment — machines nobody has rented right now. Renting that spare capacity out cheaply is a genuine win for both sides. The provider only needs to take it back the instant a full-price customer wants it. The provider earns money from otherwise-idle machines. You get a real discount, often a very large one.
The cost is real too: your job can be interrupted, without much warning, whenever that capacity is needed elsewhere.
How it works
you rent a spot GPU, at a steep discount
|
v
your job runs normally...
|
until the provider needs the machine back
|
v
you get a short warning (often under two minutes)
|
v
the machine is taken away
|
v
your job has to pick up somewhere else, or start overWhether "pick up somewhere else" or "start over" happens is entirely up to how your job was written. That is the entire subject of this lesson.
A real example you have seen
Flights, hotel rooms and cinema tickets are all sometimes sold at a discount as "standby" or "off-peak." That is spare capacity that would otherwise go completely unused. The pattern is the same: a real discount, in exchange for accepting a real risk of disruption. A spot GPU offers the same idea, applied to computing instead of travel.
The honest part
Spot GPUs are not a discount on everything. They are the right choice specifically for work that can survive an interruption. Something that saves its progress regularly, and can pick back up close to where it left off. For work that genuinely cannot be interrupted, like a live service answering real users right now, a spot GPU is a real risk. Not a free discount.
Remember this
- Spot GPUs cost far less, in exchange for the risk of being taken away with short notice.
- The provider gives a short warning, not none — enough time to react, if your code is ready to.
- They fit interruptible work well. They fit work that cannot be interrupted, poorly.
What to learn next
- CUDA, drivers and container images — making sure a new machine picking up your job actually has the right software stack ready.
- Estimating how long and how much a run will cost — weighing the real discount against the real risk of restarted work.
- Scheduling GPUs on Kubernetes — where a preempted job's replacement actually gets placed.
Developer — Code and libraries.
Setup
No installs needed — this uses only Python's standard library.
Checkpointing: what actually makes a spot instance survivable
import json
import os
import random
CHECKPOINT_FILE = "checkpoint.json"
TOTAL_STEPS = 100
CHECKPOINT_EVERY = 10
def save_checkpoint(step):
with open(CHECKPOINT_FILE, "w") as f:
json.dump({"step": step}, f)
def load_checkpoint():
if os.path.exists(CHECKPOINT_FILE):
with open(CHECKPOINT_FILE) as f:
return json.load(f)["step"]
return 0
def run_training(preempt_at=None, resume=False):
start_step = load_checkpoint() if resume else 0
for step in range(start_step, TOTAL_STEPS):
if preempt_at is not None and step == preempt_at:
print(f" [preempted at step {step}, the machine was taken away]")
return step, False
if step % CHECKPOINT_EVERY == 0 and step > start_step:
save_checkpoint(step)
save_checkpoint(TOTAL_STEPS)
return TOTAL_STEPS, True
if os.path.exists(CHECKPOINT_FILE):
os.remove(CHECKPOINT_FILE)
random.seed(0)
# Simulated: a real interruption's timing comes from the cloud provider, not
# from this script. A random step here only makes the scenario concrete.
preempt_step = random.randint(30, 70)
last_checkpoint_before_preemption = (preempt_step // CHECKPOINT_EVERY) * CHECKPOINT_EVERY
print(f"run 1: training starts fresh, will be preempted at step {preempt_step}")
run_training(preempt_at=preempt_step)
print(f" last checkpoint saved was step {load_checkpoint()}")
print("\nrun 2: a new machine picks up from the last checkpoint")
_, finished = run_training(resume=True)
print(f" finished: {finished}")
print(f" steps redone (work after the last checkpoint, lost to preemption): {preempt_step - last_checkpoint_before_preemption}")
print(f" steps saved by checkpointing (not restarted from zero): {last_checkpoint_before_preemption}")
os.remove(CHECKPOINT_FILE)run 1: training starts fresh, will be preempted at step 54 [preempted at step 54, the machine was taken away] last checkpoint saved was step 50 run 2: a new machine picks up from the last checkpoint finished: True steps redone (work after the last checkpoint, lost to preemption): 4 steps saved by checkpointing (not restarted from zero): 50
The checkpointing and resuming logic here is real and exactly how it works in practice. The interruption itself is simulated. A real provider decides when to take a machine back, not a line of Python. The specific step number will also change every time you rerun this with a different random seed. What stays true regardless: without checkpointing, this run would have lost all 54 completed steps, not just the 4 that happened after the last save.
Line-by-line walkthrough
save_checkpoint runs every CHECKPOINT_EVERY steps, not only at the very end. A checkpoint only taken at the finish line offers no protection at all. The whole point is having recent progress saved, so an interruption anywhere in the run loses only a small, bounded amount of work.
load_checkpoint() deciding where run_training starts. This is the entire recovery mechanism: on resume=True, the loop begins from the last saved step instead of zero. Nothing more complicated than that is required for the core idea.
The trade-off between CHECKPOINT_EVERY being small or large. A smaller interval loses less work per interruption. But it spends more time and I/O writing checkpoints that are usually thrown away unused. A larger interval writes less often, but risks losing more real progress when an interruption does happen.
Common mistakes
No checkpointing at all. This turns every interruption into a full restart from zero. It is the single most expensive mistake possible on spot capacity — one that can make a "cheap" spot GPU end up costing more in wasted, repeated work than a normal one would have.
Not handling the interruption warning. Cloud providers typically send a real signal (often around a two-minute warning) before reclaiming a spot machine. Code that ignores this signal loses the chance to save one last checkpoint right before being cut off. The time to do so was actually available.
Checkpointing to the machine's own local disk only. If the checkpoint is saved somewhere that disappears along with the machine, it offers no real protection. Save checkpoints somewhere that survives the machine being taken away entirely — network storage, not local disk.
Assuming spot capacity behaves the same everywhere and always. Availability and interruption frequency vary by GPU type, region, and time. A spot GPU reliably available this week can become scarce and heavily interrupted the next. Design for interruption as the normal case, not the exception.
Try it yourself
Change CHECKPOINT_EVERY from 10 to 25, and rerun with the same random seed. Compare how much work gets redone this time. Think about what that trade-off would cost in real money, for a job that actually takes hours per step.
What to learn next
- CUDA, drivers and container images — making sure a new machine picking up your job actually has the right software stack ready.
- Estimating how long and how much a run will cost — weighing the real discount against the real risk of restarted work.
- Scheduling GPUs on Kubernetes — where a preempted job's replacement actually gets placed.
Researcher — Mathematics and papers.
The economics of spot pricing
Spot pricing is typically set through a provider's internal mechanism, balancing spare capacity against demand. Historically this was often an auction-like model. More recently it is frequently a simpler "discount off on-demand price, tied to real-time spare capacity" model, across major providers. Reported discounts commonly range from 60–90% off on-demand pricing for GPU instances specifically. The exact figure varies by provider, GPU generation, region and time — always check it against current, real pricing rather than assuming it.
Interruption as a stochastic process
Modelling interruption risk formally: if interruptions arrive as a Poisson process with rate $\lambda$ (interruptions per unit time), the probability of surviving uninterrupted for at least time $t$ is $e^{-\lambda t}$, and the expected time between interruptions is $1/\lambda$. This gives a principled way to choose $\text{CHECKPOINT_EVERY}$: the checkpoint interval should be small relative to $1/\lambda$ for the specific instance type and region, not chosen arbitrarily — providers and third-party trackers publish historical interruption-rate data that can inform a real $\lambda$ estimate.
Elastic and fault-tolerant training frameworks
For distributed training specifically, naive checkpoint-and-restart wastes the entire job's progress when only one worker among many is interrupted. Elastic training frameworks — PyTorch's torchrun with elastic launch, and systems like TorchElastic and Ray Train — handle a worker being removed or added mid-job. They dynamically reconfigure the running job, rather than restarting it entirely. That is a meaningfully stronger guarantee than simple checkpointing, at real implementation complexity cost. Deep-learning-specific checkpointing systems, like DeepSpeed's asynchronous checkpointing and PyTorch Distributed Checkpoint, address a further bottleneck. Saving a large model's full state can itself take long enough to not finish before a short interruption warning expires — a real, measured problem at the scale of hundred-billion-parameter models.
Diversification as an availability strategy
Interruption risk is correlated within a single instance type, region, and availability zone. So real production spot usage typically diversifies across several interchangeable instance types and zones simultaneously, using whichever pool currently has capacity. AWS explicitly documents and supports this strategy for its Spot Fleet product, and it meaningfully reduces the probability of a total capacity loss, compared to depending on one narrow pool.
Reading
- Amazon Web Services, Amazon EC2 Spot Instances documentation — interruption notice mechanics and diversification guidance
- Or, Zhang and Freedman, AWS-CBP: Real-Time Fine-Grained Spot Instance Runtime Prediction, 2022 — modelling interruption timing rather than treating it as unpredictable
What to learn next
- CUDA, drivers and container images — making sure a new machine picking up your job actually has the right software stack ready.
- Estimating how long and how much a run will cost — weighing the real discount against the real risk of restarted work.
- Scheduling GPUs on Kubernetes — where a preempted job's replacement actually gets placed.