PyTorch Tensors

Installing PyTorch with the right CUDA build

PyTorch comes in several builds, and picking the one that matches your graphics card is the difference between using your GPU and silently ignoring it.

Read these first

On this page 5
  1. Why this exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

PyTorch comes in several versions, and you must install the one that matches your computer's hardware.

Think of buying a phone charger while travelling. The charger itself is fine — the wall socket is the problem. You need the plug that fits the socket in front of you. PyTorch is the same: one library, several builds, and you pick the build that fits your machine.

Why this exists

A GPU — graphics processing unit — is a chip originally built for video games. It can do thousands of small calculations at once, which is exactly what AI needs.

NVIDIA GPUs speak a language called CUDA. For PyTorch to use your GPU, CUDA support must be baked into the build you install.

Here is the trap. If you install the wrong build, nothing crashes. PyTorch works, trains models, and never mentions the expensive GPU sitting idle next to it. That silence is why this lesson exists.

How it works

What is in your computer?
        |
        ├─ NVIDIA graphics card      →  install a CUDA build
        ├─ Apple laptop (M-series)   →  the normal build (uses Apple's chip)
        └─ no graphics card          →  the CPU build (smaller download)

You answer one question — what hardware do I have — and the PyTorch website hands you the matching install command.

A real example you have seen

Google Colab, the free notebook site, runs PyTorch with the CUDA build already installed. That is why a model trains in minutes there and in hours on a basic laptop. Same code, same PyTorch — different build, different hardware.

Remember this

  • PyTorch is one library with several builds: CPU, CUDA, and Apple.
  • The wrong build does not crash. It quietly ignores your GPU.
  • No NVIDIA card? The CPU build is smaller and everything on this site still runs.

What to learn next

Developer — Code and libraries.

The one rule

Do not guess the command. Go to the official selector at pytorch.org/get-started/locally, click your OS, package manager and hardware, and copy what it gives you.

Everything below was verified with torch 2.5.1. The CUDA versions on offer change every few months — read the current ones off the selector.

Setup

With an NVIDIA GPU, the command looks like this (cu121 means "built for CUDA 12.1"):

bash
pip install torch --index-url https://download.pytorch.org/whl/cu121

Without one, ask for the CPU build directly:

bash
pip install torch --index-url https://download.pytorch.org/whl/cpu

Be honest with your internet plan: a CUDA build is a multi-gigabyte download (around 2.5 GB on Windows). The CPU build is roughly a tenth of that.

The two facts that clear up most confusion

You do not need to install CUDA yourself. The pip package carries its own CUDA libraries inside. The only thing your system must provide is a reasonably recent NVIDIA driver — the software that talks to the card. Installing NVIDIA's multi-gigabyte "CUDA Toolkit" is for people writing GPU code in C++, not for using PyTorch.

Plain pip install torch does not always mean GPU. On Windows, as of torch 2.5, the default package is CPU-only. On Linux the default includes CUDA. This asymmetry is the single most common reason a Windows machine with a good GPU trains on the CPU.

Check what you actually got

check_install.py
import torch
print("torch version :", torch.__version__)
print("built for CUDA:", torch.version.cuda)
print("GPU usable    :", torch.cuda.is_available())
if torch.cuda.is_available():
    print("GPU name      :", torch.cuda.get_device_name(0))
Output
torch version : 2.5.1+cu121
built for CUDA: 12.1
GPU usable    : True
GPU name      : NVIDIA RTX A6000

Your version numbers and GPU name will differ — that is expected. The line that matters is GPU usable. On a CPU build, torch.version.cuda prints None and the version string ends in +cpu.

To check your driver, run nvidia-smi in a terminal. The "CUDA Version" in its top corner is the newest CUDA your driver can serve — your PyTorch build's CUDA version must be that number or lower.

Common mistakes

is_available() is False on a machine with an NVIDIA card. You almost certainly have the CPU build. Check torch.__version__ — if it ends in +cpu, uninstall and reinstall with the --index-url line above.

Driver too old. If the build's CUDA version is newer than what nvidia-smi reports, the GPU will not be usable. Update the driver from NVIDIA — that is a much smaller download than the toolkit.

Mixing conda and pip installs. Two half-installed copies of torch shadow each other and produce import errors that look insane. Pick one tool per environment, and when in doubt, make a fresh virtual environment.

A Python version that is too new. Wheels for a fresh Python release lag by weeks. If pip says "no matching distribution found for torch", check the supported Python versions on the selector before assuming anything worse.

Try it yourself

Run check_install.py on your machine and read every line. If you have an NVIDIA card and GPU usable says False, fix it now — every later lesson gets cheaper once this works.

What to learn next

Researcher — Mathematics and papers.

The compatibility model, precisely

Three version numbers interact:

  • Driver — kernel-level, system-wide, backward compatible: a driver supports its headline CUDA version and everything older.
  • CUDA runtime — the libraries (cudart, cuBLAS, cuDNN) bundled inside the PyTorch wheel. This is what torch.version.cuda reports.
  • Compute capability — the GPU architecture generation, reported by torch.cuda.get_device_capability(), e.g. (8, 6) for Ampere consumer cards.

A wheel works when: driver's max CUDA ≥ wheel's CUDA runtime, and the wheel contains compiled kernels (or PTX for forward-compatibility) for your card's compute capability. Very new GPUs on old wheels fail with "no kernel image is available for execution on the device" — the wheel predates the architecture.

What is actually inside the wheel

The size difference between builds is almost entirely compiled GPU kernels. Each supported architecture gets its own binary code (cubins), plus PTX — an intermediate representation the driver can JIT-compile for architectures released after the wheel was built, at a first-run latency cost. torch.backends.cudnn.version() exposes the bundled cuDNN.

Non-NVIDIA hardware

  • Apple silicon: the standard macOS build ships the MPS backend (Metal Performance Shaders). Device string "mps", no separate install. Operator coverage is good but not complete; unsupported ops fall back to CPU with a warning.
  • AMD: ROCm wheels exist for Linux only, from the same selector. The device string is still "cuda" — HIP masquerades under the CUDA API, which keeps most code portable.
  • Intel GPUs: XPU support is merging into mainline as of torch 2.5+, still maturing.

Reading material

  • PyTorch NeurIPS paper: Paszke et al. (2019), PyTorch: An Imperative Style, High-Performance Deep Learning Library — the design rationale for the whole stack.
  • NVIDIA's CUDA Compatibility guide (docs.nvidia.com) — the authoritative statement of the driver/runtime rules above.

What to learn next