GPUs: Memory, Scheduling and Cost
Scheduling GPUs on Kubernetes
Kubernetes decides which physical machine each GPU workload runs on, by matching what a job asks for against what each machine actually has free.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Kubernetes decides which physical machine a GPU job runs on. It checks which machines have enough free GPUs to actually fit it.
The analogy you have already lived
A wedding hall booking manager gets requests all week. One function needs two halls, another needs one, another needs three. The manager keeps a running list of which halls are already booked. Each new request gets fitted into whatever is actually free, in whichever building has room.
If nothing free is big enough for a request, that booking has to wait. Even if the total free space across every building would technically be enough. Kubernetes is software that manages many machines as if they were one. It plays exactly this role for GPU jobs across a company's real machines.
Why it exists
A company running GPU workloads rarely has only one machine. It has many, each with a fixed number of GPUs. Every job that needs a GPU has to be placed on one specific machine. That machine must actually have enough free GPUs for it, right now.
Doing that by hand does not scale past a handful of jobs. Kubernetes automates it. You describe what a job needs, and it finds — or waits for — a machine that can actually provide it.
How it works
a new job arrives, asking for: "2 GPUs"
|
v
check every machine Kubernetes knows about
|
does any machine have 2 FREE GPUs right now?
| |
yes no
| |
place the job there the job WAITS
mark those GPUs as used ("pending"), until
something frees upThis matching happens automatically, continuously, across however many machines the company actually has.
A real example you have seen
Cloud services that let you "just deploy" a model without ever being told which physical machine it landed on. Something still had to decide that. Kubernetes' scheduler is very often the thing doing the deciding, behind the scenes, on real company infrastructure.
The honest part
This lesson's code teaches the real idea behind Kubernetes GPU scheduling: matching a request against free capacity. It uses a small Python simulation, not a real cluster. Actually running Kubernetes needs real machines, real GPUs, and real cluster software installed. None of that fits inside a single script on a laptop. What follows is honest about which part is the real mechanism. And which part is a stand-in for infrastructure you would need to see for yourself.
Remember this
- Kubernetes places each GPU job on a machine that has enough free GPUs, right now.
- A job that cannot fit anywhere currently free waits, rather than failing outright.
- The scheduling idea can be understood and practised without a real cluster — running it for real needs one.
What to learn next
- Docker for ML — packaging the workload that ends up inside a scheduled pod.
- Running on spot and preemptible GPUs — what happens when a scheduled GPU can be taken away mid-job.
- Sharing one GPU between models — the fractional-GPU sharing this lesson's whole-number scheduling model does not cover on its own.
Developer — Code and libraries.
Setup
No installs needed for the simulation — this uses only Python's standard library.
Simulating the core scheduling decision
from dataclasses import dataclass
@dataclass
class Node:
name: str
gpus_total: int
gpus_used: int = 0
@property
def gpus_free(self):
return self.gpus_total - self.gpus_used
@dataclass
class Pod:
name: str
gpus_requested: int
def schedule(nodes, pods):
"""First-fit: place each job on the first machine with enough free GPUs.
This mirrors the core idea of the Kubernetes scheduler filtering nodes
by a resource request, simplified to one resource type."""
placements, pending = {}, []
for pod in pods:
placed = False
for node in nodes:
if node.gpus_free >= pod.gpus_requested:
node.gpus_used += pod.gpus_requested
placements[pod.name] = node.name
placed = True
break
if not placed:
pending.append(pod.name)
return placements, pending
nodes = [Node("gpu-node-1", gpus_total=4), Node("gpu-node-2", gpus_total=2)]
pods = [
Pod("training-job-a", gpus_requested=2),
Pod("inference-svc-b", gpus_requested=1),
Pod("training-job-c", gpus_requested=3),
Pod("inference-svc-d", gpus_requested=1),
Pod("training-job-e", gpus_requested=2),
]
placements, pending = schedule(nodes, pods)
print("scheduling result:")
for pod_name, node_name in placements.items():
print(f" {pod_name:16} -> {node_name}")
if pending:
print("pending (no node had enough free GPUs):")
for p in pending:
print(f" {p}")
print("\nnode state after scheduling:")
for node in nodes:
print(f" {node.name}: {node.gpus_used}/{node.gpus_total} GPUs used")scheduling result: training-job-a -> gpu-node-1 inference-svc-b -> gpu-node-1 inference-svc-d -> gpu-node-1 training-job-e -> gpu-node-2 pending (no node had enough free GPUs): training-job-c node state after scheduling: gpu-node-1: 4/4 GPUs used gpu-node-2: 2/2 GPUs used
This is real, deterministic Python — rerun it and you get exactly this result every time. training-job-c asked for 3 GPUs. By the time it was considered, node-1 had only 1 free and node-2 only had 2 total. Six free GPUs existed across the whole cluster at various points. Yet this one job could not be placed — a genuine and common scheduling failure mode, not a bug in this script.
A real Kubernetes GPU request, for reference
This is correct, real Kubernetes YAML — the configuration a real job would actually use to ask for a GPU. There is no output block for it. Running it needs a real cluster with the NVIDIA device plugin installed, which cannot exist inside a single script.
apiVersion: v1
kind: Pod
metadata:
name: training-job-a
spec:
containers:
- name: trainer
image: my-training-image:latest
resources:
limits:
nvidia.com/gpu: 2 # this pod is asking for exactly 2 GPUsnvidia.com/gpu is a special resource name the NVIDIA device plugin registers with Kubernetes. The scheduler can then count and allocate whole GPUs, the same way it already counts CPU and memory.
Line-by-line walkthrough
first-fit, checking nodes in order. Real Kubernetes uses a more sophisticated scoring system by default. But the underlying filter — "does this node have enough of what was asked for?" — is the same idea this simulation makes concrete.
training-job-c going pending. This is the single most important thing to notice in the whole demo. Pending is not an error. It means the scheduler correctly found nowhere the job currently fits, and is waiting for room, exactly as designed.
GPUs are requested as a whole number, never a fraction, in standard Kubernetes. nvidia.com/gpu: 2 means exactly two entire GPUs, allocated to this one pod alone. Splitting a single GPU between pods needs one of the sharing mechanisms from sharing one GPU between models, not a fractional request here.
Common mistakes
Requesting more GPUs than any single machine in the cluster has. A request for 8 GPUs will sit pending forever on a cluster where the biggest machine only has 4. That holds even if the cluster's total GPU count across many machines is much higher than 8 — Kubernetes schedules onto one node, not split across several, for a single pod.
Assuming "pending" means something is broken. As the simulation shows, a job can be perfectly correctly configured and still wait, because of real, legitimate resource pressure. Check node capacity before assuming a configuration problem.
Fragmentation from placing large jobs after small ones. In the simulation, placing the 3-GPU job earlier would have let it schedule successfully — before node-1 filled up with smaller ones. Real schedulers offer bin-packing strategies specifically to reduce this kind of waste; see the researcher section.
Forgetting the device plugin. A cluster with physical GPUs will not report nvidia.com/gpu as a schedulable resource at all. Not until the NVIDIA device plugin, or an equivalent, is installed and running. Without it, GPU requests are invalid outright, not only unscheduled.
Try it yourself
Reorder the pods list so training-job-c (3 GPUs) is scheduled first, before the smaller jobs. Rerun the simulation and see whether every job gets placed this time.
What to learn next
- Docker for ML — packaging the workload that ends up inside a scheduled pod.
- Running on spot and preemptible GPUs — what happens when a scheduled GPU can be taken away mid-job.
- Sharing one GPU between models — the fractional-GPU sharing this lesson's whole-number scheduling model does not cover on its own.
Researcher — Mathematics and papers.
Bin-packing and scheduler scoring
Kubernetes' default scheduler operates in two phases. Filtering asks which nodes even have enough of every requested resource — CPU, memory, GPU, any other declared resource. Scoring then asks, among the nodes that pass filtering, which is the "best" placement, by a configurable set of priority functions. The default scoring includes strategies like LeastAllocated — spread load across nodes, leaving more room for future large requests — and MostAllocated: pack tightly, freeing whole nodes for other uses or for scale-down. This is a direct, tunable answer to the fragmentation problem the demo's simulation deliberately exposes.
For GPU workloads specifically, packing quality matters more than for many other resources. GPUs are typically requested in whole units, and are individually expensive. A cluster with plenty of aggregate free GPU capacity, spread as single free GPUs across many nodes, cannot satisfy a job requesting several GPUs on one node — exactly the failure mode training-job-c hits in the demo.
The NVIDIA device plugin and resource extension
Kubernetes has no built-in concept of a GPU. The device plugin framework, a stable Kubernetes extension point since 1.10, lets a vendor register custom schedulable resources. The NVIDIA device plugin discovers physical GPUs on each node via NVML, and advertises them as the nvidia.com/gpu resource — making the earlier YAML manifest meaningful to the scheduler. MIG-partitioned GPUs (see sharing one GPU between models) are advertised as additional, distinctly-named resources per partition size, letting the scheduler place jobs onto specific MIG slices directly.
Beyond the default scheduler
For GPU-heavy, ML-specific workloads, several purpose-built schedulers extend or replace Kubernetes' default:
- Kubeflow / Volcano — add gang scheduling and queueing with priority. Gang scheduling means all-or-nothing placement, for distributed training jobs that need every worker to start together, which the default scheduler does not guarantee.
- Run:ai / KAI Scheduler — add GPU fractionalisation, fair-share quotas across teams, and bin-packing strategies tuned specifically for GPU fragmentation.
These exist because the default scheduler's per-pod, independent scheduling model is a poor fit for distributed training jobs. Their workers must all be scheduled together to make progress — a gap gang scheduling closes directly.
Reading
- Kubernetes documentation, Scheduling, Preemption and Eviction — the filtering/scoring model in full
- NVIDIA, GPU Operator and k8s-device-plugin documentation — the concrete mechanism behind
nvidia.com/gpu - Xiao et al., Gandiva: Introspective Cluster Scheduling for Deep Learning, OSDI 2018 — an influential design for GPU-cluster-aware scheduling beyond generic Kubernetes defaults
What to learn next
- Docker for ML — packaging the workload that ends up inside a scheduled pod.
- Running on spot and preemptible GPUs — what happens when a scheduled GPU can be taken away mid-job.
- Sharing one GPU between models — the fractional-GPU sharing this lesson's whole-number scheduling model does not cover on its own.