GPUs: Memory, Scheduling and Cost

CUDA, drivers and container images

A GPU program needs its software layers to actually match the hardware and each other. A mismatch does not always show an obvious error — sometimes it silently refuses to use the GPU at all.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A GPU program needs several software layers to actually match the hardware and each other. Otherwise it quietly fails to use the GPU at all.

The analogy you have already lived

An appliance bought in one country often will not plug into a wall socket in another. The plug shape might not fit. Or worse, the voltage might be wrong, even if you force an adapter onto it. Getting this wrong does not always announce itself loudly. Sometimes the appliance does not turn on at all, or works badly, with no obvious explanation.

Running code on a GPU has exactly this kind of layered compatibility requirement. Several "sockets" all have to match at once.

Why it exists

Using a GPU is not just "the code" and "the GPU." There is a real stack of software layers between them. Every layer has to actually agree with the ones next to it:

  • The driver — software installed on the machine itself, which lets the operating system talk to the physical GPU at all.
  • CUDA — NVIDIA's toolkit that lets programs like PyTorch send work to the GPU, at a specific version.
  • The framework (PyTorch, TensorFlow) — built against one specific CUDA version.
  • The container image, if you use one. It bundles its own CUDA version inside it, separate from whatever is installed on the actual machine.

Any one of these being the wrong version relative to the others can mean the GPU is never actually used. Even though every individual piece "installed successfully."

How it works

   physical GPU
       |
   NVIDIA DRIVER            (installed on the machine itself)
       |
   CUDA TOOLKIT VERSION      (must be supported by the driver)
       |
   PYTORCH / TENSORFLOW      (built against one specific CUDA version)
       |
   YOUR CODE

Each arrow is a real compatibility requirement. A newer driver usually supports older CUDA versions built against it — that direction tends to be forgiving. Expecting an old driver to run a program built for a much newer CUDA version is where mismatches actually bite.

A real example you have seen

A phone app that refuses to install because your phone's operating system is "too old" for it. The app was built assuming a newer version of something underneath it. GPU software has the same layered dependency, with more layers involved. The failure is sometimes far quieter than an app store outright refusing to install.

The honest part

The most frustrating version of this mismatch is not a clear error message. It is code that runs, produces no crash, and runs on the CPU instead of the GPU. Silently, without complaint, often much slower — with nothing visibly wrong until someone notices performance is far worse than expected.

Remember this

  • Using a GPU depends on a stack of matching software versions, not just the GPU itself.
  • A driver usually supports older CUDA versions built against it, not always newer ones.
  • The worst version of this mismatch is silent — code that quietly runs on the CPU instead of failing loudly.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

Checking your own stack, for real

check_gpu_stack.py
import subprocess
import torch

driver_out = subprocess.run(
    ["nvidia-smi", "--query-gpu=driver_version,name", "--format=csv,noheader"],
    capture_output=True, text=True, timeout=5,
).stdout.strip()
driver_version, gpu_name = [s.strip() for s in driver_out.split(",")]

print(f"GPU:                        {gpu_name}")
print(f"NVIDIA driver version:      {driver_version}")
print(f"PyTorch was built for CUDA: {torch.version.cuda}")
print(f"PyTorch version:            {torch.__version__}")
print(f"CUDA available to PyTorch:  {torch.cuda.is_available()}")

if torch.cuda.is_available():
    x = torch.randn(3, 3, device="cuda")
    y = x @ x
    print(f"a real matrix multiply on the GPU worked: result shape {tuple(y.shape)}")
Output
GPU:                        NVIDIA RTX A6000
NVIDIA driver version:      556.39
PyTorch was built for CUDA: 12.1
PyTorch version:            2.5.1+cu121
CUDA available to PyTorch:  True
a real matrix multiply on the GPU worked: result shape (3, 3)

Real output from this machine's actual GPU stack. Yours will show different version numbers, and that is completely fine. Check not for any specific number, but that CUDA available to PyTorch reads True, and the matrix multiply line actually appears — confirming the GPU genuinely did the work.

Line-by-line walkthrough

nvidia-smi --query-gpu=driver_version. Asks the driver directly for its own version. That is the ground truth for what the machine actually has installed, independent of anything Python or PyTorch reports.

torch.version.cuda. This is the CUDA version PyTorch was built against when it was compiled. Not the CUDA version installed on your machine. A common point of confusion: this number can differ from what nvcc --version reports on the machine itself. That is expected, not a bug.

torch.cuda.is_available(). The single most important line for catching a mismatch. Say this is False on a machine that visibly has a working GPU, visible to nvidia-smi. The driver and the installed PyTorch build do not agree. The reason is almost always a version mismatch somewhere in the stack.

A matching Docker base image, for reference

Dockerfile
# The CUDA version in this base image tag must be supported by the HOST
# machine's driver -- see the compatibility table below.
FROM nvidia/cuda:12.1.0-runtime-ubuntu22.04

RUN pip install torch --index-url https://download.pytorch.org/whl/cu121

There is no output block for this. Building and running a GPU-enabled container needs the NVIDIA Container Toolkit installed on the host, which cannot be demonstrated inside a single script. The line to focus on is the version match between the base image's CUDA tag (12.1.0) and the PyTorch build's expected CUDA version (cu121). Both must also be supported by the host machine's actual driver.

Common mistakes

Assuming a container fixes every compatibility problem. A container bundles its own CUDA toolkit, but it still needs the host machine's driver to support that CUDA version. Containerising code does not remove the driver requirement. It only removes the CUDA toolkit and framework install from the list of things that can mismatch.

Checking torch.__version__ and stopping there. The framework version alone says nothing about whether it can actually reach the GPU. torch.cuda.is_available() is the check that actually matters.

Silently falling back to CPU without noticing. Code that never explicitly checks torch.cuda.is_available() before running can execute correctly on the CPU for a long time, before anyone notices it never touched the GPU at all. Correct answers, at a small fraction of the expected speed.

Mixing driver versions across a fleet of machines. A model tested and confirmed working on one machine can silently fail to use the GPU on another machine in the same fleet. That happens if that machine's driver was never updated to match. Check this per machine, not once for the whole fleet.

Try it yourself

Run nvcc --version in a terminal, if it is installed, and compare its reported CUDA version against torch.version.cuda. They are allowed to differ — understanding why they are allowed to differ is the point of the exercise.

What to learn next

Researcher — Mathematics and papers.

The driver–toolkit compatibility model

NVIDIA's compatibility model is built around backward compatibility within a major version series. A given driver version supports a maximum CUDA toolkit version, published in NVIDIA's official CUDA compatibility matrix, and generally supports all CUDA minor versions at or below that maximum. This is why a fresh driver install tends to be forgiving. The moment of real risk is an old driver on a machine that has not been updated, encountering a framework built against a CUDA version newer than that driver supports.

CUDA Forward Compatibility packages exist specifically to relax this constraint for datacenter GPUs, letting a slightly older driver run against a newer CUDA toolkit than it would otherwise support — relevant for large, hard-to-update fleets where updating every driver in lockstep with every framework release is operationally expensive.

Two separate version numbers that both matter

nvidia-smi's reported "CUDA Version" in its header is the maximum CUDA version the installed driver supports — not the CUDA toolkit actually installed on the machine, which is a separate, independently versioned thing (queried via nvcc --version when the toolkit itself is present). Confusing these two numbers is one of the most common sources of "but it says CUDA 12.5 right there" debugging dead ends. The driver supporting a version is necessary, but not sufficient, for a program built against it to actually run.

Container image layering in practice

NVIDIA's official CUDA container images come in three tiers, each bundling progressively more: base (CUDA runtime libraries only), runtime (adds cuDNN and other runtime libraries typical ML frameworks need), and devel (adds the full compiler toolchain, for building CUDA code from source inside the container). Choosing the smallest tier that satisfies your actual dependencies keeps image size and attack surface down. A common and avoidable mistake is shipping a devel image to production when runtime would have worked, purely because it was the tag used during development.

The NVIDIA Container Toolkit is the component that makes any of this work at all in a containerised setting — it is what exposes the host's GPU devices and driver libraries into an otherwise fully isolated container namespace at docker run --gpus all time, bridging the host driver to the container's CUDA toolkit rather than requiring a driver be installed inside the container image itself, which is neither supported nor necessary.

Reading

  • NVIDIA, CUDA Compatibility documentation — the official driver/toolkit compatibility matrix and forward-compatibility packages
  • NVIDIA, NVIDIA Container Toolkit documentation — the mechanism bridging host driver to container CUDA runtime

What to learn next

What to learn next

These follow on from what you just read.

  • GPUs: Memory, Scheduling and Cost

    Serverless GPU platforms

    Serverless GPU platforms give you a GPU only for the seconds you actually use one, and hand it back afterwards — in exchange for a real, measurable cold-start cost.

  • GPUs: Memory, Scheduling and Cost

    GPU cost per request

    The true cost of one request is the GPU's price divided by how many requests it actually serves — and that number depends far more on utilisation than on the GPU's hourly price.

  • GPUs: Memory, Scheduling and Cost

    When a CPU is enough

    A GPU is not free, and not every model needs one. For a small model with modest traffic and a forgiving latency budget, the CPU you already have can be the right, cheaper answer.