When the GPU is slower than the CPU
A GPU wins through parallel bulk work — for small tensors, chatty little operations, or anything dominated by transfer and launch overhead, the CPU honestly wins.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A GPU is a cargo truck: unbeatable for a full load, absurd for delivering one letter.
Sending a letter across town by bicycle takes twenty minutes. Sending it by cargo truck takes longer — the truck needs loading, paperwork, a driver, unloading. The truck wins only when there are ten thousand letters. Nobody calls the truck slow. It was built for bulk.
A GPU is that truck. Thousands of workers that shine when there is enough work for all of them at once.
Why this surprises people
"GPU = fast" gets repeated so often that the fine print disappears. The fine print is fixed costs. Every piece of GPU work pays them:
- The ferry. Data must travel to the GPU's separate memory and back — the transfer cost.
- The paperwork. Each operation sent to the GPU costs a small fixed fee to launch, whatever its size.
For big work, these costs vanish into the total. For tiny work, they are the total. A tiny model on a tiny batch can run several times faster on a plain CPU.
How it works
work size: tiny medium huge
CPU time: instant okay slow and slower
GPU time: overhead overhead overhead + fast bulk work
winner: CPU depends GPU (by miles)There is a crossover point. Below it, the CPU wins honestly. Above it, the GPU pulls away and never looks back.
A real example you have seen
Opening a calculator app to add 7 and 5 versus doing it in your head. The app computes faster than any human — but unlocking the phone and opening the app is the overhead. For one small sum, your head wins. For a tax return, the app does.
Remember this
- GPUs win through parallel bulk work, not through being fast at everything.
- Transfer and launch overhead are fixed taxes; tiny work drowns in them.
- Measure before assuming — the crossover is real and closer than people think.
What to learn next
- Timing GPU code correctly — the measurement discipline behind every claim here.
- DataLoader workers and speed — keeping the truck loaded.
- DataParallel vs DistributedDataParallel — when one truck is no longer enough.
Developer — Code and libraries.
Setup
pip install torchThis lesson needs a machine with both processors to race them. Numbers captured with torch 2.5.1 on an NVIDIA RTX A6000 and its host CPU; your crossover point will sit elsewhere, but there will be one.
Race them across sizes
import torch
import time
if not torch.cuda.is_available():
raise SystemExit("needs a GPU - the point is the comparison")
def bench(device, n):
a = torch.randn(n, n, device=device)
b = torch.randn(n, n, device=device)
a @ b # warm-up
if device == "cuda":
torch.cuda.synchronize()
start = time.perf_counter()
for _ in range(10):
a @ b
if device == "cuda":
torch.cuda.synchronize()
return (time.perf_counter() - start) / 10 * 1000
print(f"{'size':>10} {'cpu ms':>10} {'gpu ms':>10} winner")
for n in (32, 128, 512, 2048):
cpu, gpu = bench("cpu", n), bench("cuda", n)
print(f"{n:>10} {cpu:>10.3f} {gpu:>10.3f} {'GPU' if gpu < cpu else 'CPU'}") size cpu ms gpu ms winner
32 0.005 0.019 CPU
128 0.029 0.071 CPU
512 0.431 0.041 GPU
2048 19.064 0.816 GPUAt size 32 the CPU is four times faster. At 2048 the GPU is twenty-three times faster. Both facts are true; neither is a malfunction. Notice the GPU column barely moves from 32 to 512 — that flat region is the launch overhead, the paperwork fee that dominates until the work grows past it. All timing rules from the timing lesson apply — without the synchronize() calls this table would be fiction.
The ferry fee, isolated
import torch
import time
if not torch.cuda.is_available():
raise SystemExit("needs a GPU")
x = torch.randn(64, 64)
torch.relu(x.cuda()).cpu() # warm-up: CUDA startup is paid here, once
torch.cuda.synchronize()
def best_of(fn, repeats=20):
times = []
for _ in range(repeats):
start = time.perf_counter()
fn()
times.append(time.perf_counter() - start)
return min(times) * 1000
cpu_ms = best_of(lambda: torch.relu(x))
trip_ms = best_of(lambda: torch.relu(x.cuda()).cpu())
print(f"relu on CPU: {cpu_ms:8.4f} ms")
print(f"ship + relu on GPU + return: {trip_ms:8.4f} ms")
print(f"the ferry cost {trip_ms / cpu_ms:.0f}x more than the work")relu on CPU: 0.0014 ms ship + relu on GPU + return: 0.1385 ms the ferry cost 97x more than the work
The operation itself is tiny on both processors. Ninety-seven times the cost is the round trip. This is why "move to GPU, do one small thing, move back" — inside a loop — is the deadliest pattern in beginner GPU code.
When the CPU genuinely wins
- Small models on small batches — the sweep's left column, common in classical ML and tiny MLPs.
- Inference for one sample at a time — batch size 1 rarely feeds a GPU; see batching and throughput thinking before buying GPU serving.
- Chatty code — loops of tiny ops, per-element Python, frequent
.item()calls. Each is a paperwork fee plus, often, a forced wait. - Preprocessing — tokenising, parsing, image decoding are CPU jobs; that is why DataLoader workers exist.
Common mistakes
Fixing the verdict, not the workload. The GPU losing on tiny batches is information: batch more work together, fuse ops with torch.compile, or accept the CPU. Buying a bigger GPU does not shrink the overhead.
Round-tripping for convenience. Calling .cpu().numpy() mid-loop for logging or a NumPy helper drags the ferry fee — and a forced synchronisation — into every step. Accumulate on-device; ferry once per epoch.
Benchmarking without synchronisation and declaring victory. The unsynchronised GPU always "wins". See timing GPU code correctly before trusting any table, including this one.
Assuming Colab's GPU makes everything faster. Switching runtime type speeds up nothing by itself; small workloads get slower. Measure your actual case.
Try it yourself
Extend the sweep with sizes 64, 256 and 1024 and find your machine's crossover size. Then rerun the whole sweep with dtype=torch.float16 tensors and watch the crossover move — the GPU's bulk advantage grows, the overhead stays.
What to learn next
- Timing GPU code correctly — the measurement discipline behind every claim here.
- DataLoader workers and speed — keeping the truck loaded.
- DataParallel vs DistributedDataParallel — when one truck is no longer enough.
Researcher — Mathematics and papers.
A two-parameter cost model
Model an offloaded operation as:
$$ T_{\text{gpu}}(n) = \alpha + \frac{W(n)}{\beta_{\text{gpu}}}, \qquad T_{\text{cpu}}(n) = \frac{W(n)}{\beta_{\text{cpu}}} $$
- $\alpha$ — fixed overhead per offload: kernel launch (~5–10 µs each) plus any PCIe transfer (latency ~10 µs, bandwidth ~16–32 GB/s on PCIe 4/5).
- $W(n)$ — work, e.g. $2n^3$ FLOPs for an $n \times n$ matmul.
- $\beta$ — sustained throughput of each processor for this op.
The crossover satisfies $W(n^*) = \alpha \cdot \left( \beta_{\text{gpu}}^{-1} - \beta_{\text{cpu}}^{-1} \right)^{-1} \approx \alpha \beta_{\text{gpu}}$ when $\beta_{\text{gpu}} \gg \beta_{\text{cpu}}$: the GPU must be handed at least $\alpha \beta_{\text{gpu}}$ worth of work per launch to break even. With $\alpha \approx 10$ µs and $\beta_{\text{gpu}} \approx 10^{13}$ FLOP/s, that is $10^8$ FLOPs per launch — a 370×370 matmul, matching the sweep's observed crossover between 128 and 512.
The same inequality in memory-bound terms uses arithmetic intensity and the roofline model: ops below the ridge point are bandwidth-limited on both processors, and the GPU's advantage shrinks to the bandwidth ratio (~10x) rather than the FLOP ratio (~100x) — before overhead.
Beyond the model: second-order effects
- Occupancy: a kernel using 2 of 84 SMs pays full launch cost for 2% of the machine. Effective $\beta_{\text{gpu}}$ is workload-dependent, not a datasheet constant.
- Amortisation strategies: CUDA Graphs replay whole launch sequences with one submission, attacking $\alpha$ directly; torch.compile reduces the number of launches by fusion. Both move the crossover left without touching hardware.
- Dispatch overhead: for microsecond-scale ops, PyTorch's Python and dispatcher overhead (~1–5 µs per op) rivals the launch itself — one reason CPU eager mode wins tiny workloads by even more than the model predicts.
- Unified-memory architectures (Apple silicon, integrated SoCs) delete the PCIe term but keep launch overhead and the occupancy question; the crossover survives, smaller.
References
- Williams, Waterman, Patterson (2009), Roofline: an insightful visual performance model for multicore architectures — the bandwidth-vs-compute lens.
- NVIDIA, CUDA C++ Best Practices Guide — launch overhead, transfer amortisation, and the "batch small transfers" doctrine.
- Gregg and Hazelwood (2011), Where is the data? Why you cannot debate CPU vs. GPU performance without the answer — the transfer-inclusive benchmarking argument this lesson demonstrates.
What to learn next
- Timing GPU code correctly — the measurement discipline behind every claim here.
- DataLoader workers and speed — keeping the truck loaded.
- DataParallel vs DistributedDataParallel — when one truck is no longer enough.