Moving tensors between CPU and GPU
Tensors live on a device, operations demand all their tensors on the same one, and .to(device) is the moving van — with a fare you should know about.
- 7 min read
- 3 reading levels
- Published
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Every tensor lives on one device — the CPU or a GPU — and two tensors can only work together if they live on the same one.
Think of two kitchens in one restaurant: the small one where ingredients arrive, and the huge one where cooking is fast. A dish must be prepared in one kitchen. Half the ingredients upstairs and half downstairs means someone has to carry things over before any cooking happens.
The CPU is the small kitchen — where every tensor is born. The GPU is the huge one — thousands of hands, brilliant at repetitive work.
Why it exists
The two processors have separate memory. The GPU cannot casually reach into CPU memory; numbers must be shipped across, and shipping takes real time.
PyTorch refuses to hide this cost from you. You move tensors on purpose, and mixing kitchens raises an error rather than triggering a slow secret transfer. Annoying at first, and then you appreciate it: the errors are cheaper than the invisible slowness would be.
How it works
CPU memory GPU memory
[ tensor ] --- .to("cuda") ---> [ tensor ] fast maths happens here
[ result ] <--- .cpu() -------- [ result ]
rule: one operation = all tensors on one deviceThe pattern of nearly every training program: ship the model over once, ship each batch of data as it is needed, do the heavy maths there, bring small results back.
A real example you have seen
Free GPUs on Google Colab make this visible. The same tiny model trains noticeably faster after "Change runtime type → GPU" — but only if the code moved its tensors. Forgetting the move is the classic Colab experience: paying for a Ferrari and pushing it.
Remember this
- Tensors are born on the CPU; moving them is your explicit job.
- One operation needs all of its tensors on one device.
- Moving costs time — move rarely, compute plenty in between.
What to learn next
- Seeds and reproducible runs — the other global state your experiments depend on.
- Installing PyTorch with the right CUDA build — if
cuda.is_available()disagreed with your hardware. - PyTorch basics — the training loop this pattern slots into.
Developer — Code and libraries.
Setup
pip install torchThis lesson runs correctly with or without a GPU. Outputs below were captured with torch 2.5.1 on a machine with an NVIDIA card; the noted lines differ on CPU-only machines.
The portable pattern
import torch
device = "cuda" if torch.cuda.is_available() else "cpu"
print("using:", device)
x = torch.ones(3) # tensors are born on the CPU
print(x.device)
x = x.to(device) # .to() COPIES it to the device
print(x.device)
y = torch.ones(3, device=device) # or create it there directly
z = x + y # both operands on one device: allowed
print(z)
back = z.cpu().numpy() # NumPy only understands CPU memory
print(back)using: cuda cpu cuda:0 tensor([2., 2., 2.], device='cuda:0') [2. 2. 2.]
On a machine without a GPU, the same file prints using: cpu, every device reads cpu, and no line crashes. That is the point of the first line — write it once at the top and the code runs everywhere.
The walkthrough
x = x.to(device) — note the assignment. .to() returns a new tensor on the target device; it does not move the original. x.to(device) alone, without the =, is the most common single-line GPU bug in existence: it computes a copy and throws it away.
cuda:0 means "the first GPU" — the :0 is an index, visible when a machine has several.
Mixing devices fails loudly:
a = torch.ones(3)
b = torch.ones(3, device="cuda")
print(a + b)RuntimeError: Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu!
You will meet this message many times. It nearly always means one tensor missed the ferry — often the labels while the inputs made it, or a freshly created tensor inside the loop.
z.cpu().numpy() — the return trip. NumPy, matplotlib, pandas and printing large results all need CPU data. Calling .numpy() on a CUDA tensor raises an error telling you to call .cpu() first.
Models move the same way: model.to(device) moves every parameter inside (in place, unlike tensors — one of PyTorch's stranger asymmetries). Then each batch moves in the loop: inputs = inputs.to(device).
Common mistakes
Moving inside the hot loop unnecessarily. Shipping the model every batch, or bouncing results back per step for printing, drowns the GPU's speed in ferry rides. Move the model once; accumulate metrics on-device and bring one number back per epoch.
Timing without waiting. GPU work runs asynchronously — Python queues the work and races ahead. Timing with time.time() right after an operation measures the queueing, not the work. Call torch.cuda.synchronize() before reading the clock.
Assuming the GPU is always faster. For tiny tensors, transfer plus launch overhead exceeds the compute. A GPU earns its keep on big tensors and big models; a 10-element regression is faster at home on the CPU.
Hardcoding "cuda". Code with a literal "cuda" inside crashes on every machine without one — including your reader's. The device = ... if ... else line at the top costs nothing and makes the file portable.
Loading GPU-saved checkpoints on a CPU box. torch.load(path) tries to restore tensors to their original device and fails on machines without one. Pass map_location="cpu".
Try it yourself
If you have a GPU: create two (4000, 4000) random tensors and matrix-multiply them on CPU, then on GPU, timing both — with a synchronize() before each clock read. Then repeat at (50, 50) and watch the ranking flip. No GPU? Predict both outcomes, and keep the portable pattern.
What to learn next
- Seeds and reproducible runs — the other global state your experiments depend on.
- Installing PyTorch with the right CUDA build — if
cuda.is_available()disagreed with your hardware. - PyTorch basics — the training loop this pattern slots into.
Researcher — Mathematics and papers.
The transfer is the tax
Host-to-device copies traverse PCIe: ~16 GB/s effective on PCIe 4.0 x16, ~32 GB/s on 5.0 — versus 1000+ GB/s HBM bandwidth on the device itself. A float32 batch of shape (256, 3, 224, 224) is ~154 MB, or ~10 ms of PCIe 4.0 time; a training step that computes in 20 ms and ships 10 ms of data has capped its own utilisation. This ratio drives the DataLoader design of overlapping transfer with compute.
Two mechanisms enable overlap:
- Pinned (page-locked) host memory —
tensor.pin_memory(), orpin_memory=Truein a DataLoader. DMA engines require pinned pages; pageable memory forces an extra staging copy. - Asynchronous copies —
x.to(device, non_blocking=True)from pinned memory enqueues the copy on a CUDA stream and returns immediately, letting the previous batch's compute proceed concurrently.
Without both together, non_blocking=True silently degrades to a synchronous copy — it is an optimisation hint, not a guarantee.
The asynchronous execution model
CUDA operations are enqueued to streams (ordered queues) and execute after the Python call returns. Errors can therefore surface at a later, unrelated line — the origin of the advice to re-run with CUDA_LAUNCH_BLOCKING=1 (or on CPU) for honest stack traces. Synchronisation points — .item(), .cpu(), printing a CUDA tensor, boolean mask indexing — force the host to wait, which is why a stray .item() per step in a hot loop is a real performance bug and why metric accumulation should stay on-device.
Topology beyond one GPU
Multi-GPU systems add device-to-device paths: peer-to-peer over PCIe, NVLink (hundreds of GB/s), and collectives via NCCL. The programming model stays the same — tensors pinned to cuda:k, explicit movement — with torch.distributed layering allreduce-style communication on top. On unified-memory architectures (Apple silicon MPS, integrated SoCs) the copy disappears physically but the API keeps the same explicit device semantics for portability.
References
- Paszke et al. (2019), PyTorch: An Imperative Style, High-Performance Deep Learning Library, NeurIPS — §"memory management" and the CUDA caching allocator, which is also why
nvidia-smishows more memory used than your tensors need: freed blocks are cached for reuse, not returned to the driver. - NVIDIA, CUDA C++ Best Practices Guide — the authoritative treatment of pinned memory, streams and transfer overlap.
What to learn next
- Seeds and reproducible runs — the other global state your experiments depend on.
- Installing PyTorch with the right CUDA build — if
cuda.is_available()disagreed with your hardware. - PyTorch basics — the training loop this pattern slots into.