GPU Memory and Speed

Timing GPU code correctly

The GPU works asynchronously, so a normal stopwatch measures only the order-taking — synchronise before reading the clock, or use CUDA events, or your benchmark is fiction.

On this page 5
  1. Why it works this way
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

GPU work happens after your Python line finishes, so timing the line measures almost nothing.

Think of ordering at a busy dhaba. You say "two thalis" and the waiter nods in one second. The food arrives fifteen minutes later. If you time the nod and announce "this dhaba serves food in one second", you have measured the order-taking, not the cooking.

Python is you; the GPU is the kitchen. Nearly every GPU command returns as soon as the order is queued.

Why it works this way

This design is on purpose, and it is a good one. While the GPU cooks one batch, Python can run ahead — preparing the next batch, queueing more work. Neither side waits for the other. Training would be much slower if every order blocked until served.

The price is that a naive stopwatch lies to you. The lie flows in a painful direction: it makes GPU code look absurdly fast. People publish speed comparisons where the GPU "wins" because it was never actually timed.

How it works

wrong:
  start clock -> give order -> stop clock        (measures the nod)
                      |
                      v  GPU still cooking...

right:
  start clock -> give order -> WAIT until served -> stop clock
                                 (synchronise)

The waiting command is called synchronise: it makes Python stand at the counter until the GPU finishes everything queued so far.

A real example you have seen

Anyone who has clicked "download" and watched the button turn green instantly knows this. The button acknowledged you; the file is still coming. Judging download speed by the button is the same mistake as timing a GPU with a normal clock.

Remember this

  • GPU commands return when queued, not when done.
  • Always synchronise before reading the clock, at both ends.
  • Warm up before timing — first calls pay one-time setup costs.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

This lesson's point only exists on a GPU. Outputs captured with torch 2.5.1 on an NVIDIA RTX A6000; your numbers will differ, the trap is identical everywhere.

The lie, measured

timing_wrong.py
import torch
import time

if not torch.cuda.is_available():
    raise SystemExit("needs a GPU - the async behaviour IS the lesson")

a = torch.randn(4096, 4096, device="cuda")
b = torch.randn(4096, 4096, device="cuda")
a @ b                                  # warm-up: first CUDA calls pay setup costs
torch.cuda.synchronize()

start = time.perf_counter()
c = a @ b
wrong = (time.perf_counter() - start) * 1000
torch.cuda.synchronize()               # let that work finish before round two

start = time.perf_counter()
c = a @ b
torch.cuda.synchronize()               # wait for the GPU to actually finish
right = (time.perf_counter() - start) * 1000

print(f"without synchronize: {wrong:7.3f} ms   <- only measured the queueing")
print(f"with synchronize:    {right:7.3f} ms   <- the real duration")
Output
without synchronize:   0.124 ms   <- only measured the queueing
with synchronize:      8.586 ms   <- the real duration

A 70x lie. The unsynchronised number is the cost of enqueueing a matrix multiply; the real work took 8.6 ms. Any benchmark you have seen without a synchronize() deserves suspicion.

The cleaner tool: CUDA events

synchronize() inside a loop stalls the pipeline you are trying to measure. CUDA events are timestamps recorded by the GPU itself, in the queue, so the pipeline flows freely:

timing_events.py
import torch

if not torch.cuda.is_available():
    raise SystemExit("needs a GPU")

a = torch.randn(4096, 4096, device="cuda")
b = torch.randn(4096, 4096, device="cuda")
a @ b                                   # warm-up

start = torch.cuda.Event(enable_timing=True)
end = torch.cuda.Event(enable_timing=True)

start.record()                          # a stopwatch that lives on the GPU
for _ in range(10):
    a = a @ b
end.record()
torch.cuda.synchronize()                # wait so elapsed_time can be read

print(f"10 matmuls took {start.elapsed_time(end):.1f} ms on the GPU clock")
Output
10 matmuls took 62.7 ms on the GPU clock

The events ride the queue with the work. One synchronise at the very end, only so the host may read the result.

The benchmarking checklist

  1. Warm up — first calls include CUDA context setup, allocator growth, cuDNN algorithm search, and torch.compile compilation if enabled.
  2. Synchronise before starting the clock — else you time leftovers from earlier lines.
  3. Repeat and aggregate — take the median or minimum of many runs, not a single sample.
  4. Use real shapes — kernel selection depends on sizes; toy shapes measure different kernels.

Common mistakes

Hidden synchronisations doing you a "favour". .item(), .cpu(), print(tensor) and tensor.tolist() all force a wait. A benchmark with a stray print inside times correctly by accident — then someone removes the print and the numbers "improve". Nothing got faster.

Timing across a device transfer without sync. x.to("cuda", non_blocking=True) can return before the copy lands. Measuring transfer speed needs the same discipline as measuring compute.

Comparing a warm function against a cold one. Classic when comparing eager vs compiled, or CPU vs GPU: one side paid its setup inside the timer. Warm both, then race them.

Benchmarking with the profiler on. The profiler adds overhead. Use it to find the bottleneck; use clean timing to confirm the fix.

Try it yourself

Delete the middle torch.cuda.synchronize() in timing_wrong.py and rerun. The "with synchronize" number inflates — it now also waits for the previous, unfinished matmul. One missing line, and even the honest measurement turns dishonest.

What to learn next

Researcher — Mathematics and papers.

The execution model underneath

CUDA operations are enqueued onto streams — FIFO queues consumed by the device. The default stream serialises kernels launched onto it; the host thread continues immediately after enqueue (launch latency ~5–10 µs). Errors surface asynchronously, which is why a crash's Python stack trace can point at an innocent later line, and why CUDA_LAUNCH_BLOCKING=1 (serialise everything) is a debugging tool and a benchmarking catastrophe.

cudaEvent timestamps are recorded by the GPU when the preceding work in the stream completes, with sub-microsecond resolution; elapsed_time is device-clock arithmetic, immune to host jitter, scheduler noise and Python overhead. For multi-stream workloads, events must be recorded on the stream doing the work — an event on stream 0 does not bound kernels on stream 7.

Statistics of benchmarks

GPU clocks are not constant: DVFS ramps frequency with load and temperature, so early iterations run at boost clocks a sustained run cannot hold, and a thermally saturated card runs slower than a cold one. Consequences:

  • Report the median over many iterations after a sustained warm-up; the minimum estimates the noise floor, the mean is polluted by stragglers.
  • Fix clocks (nvidia-smi -lgc) for kernel-vs-kernel comparisons when possible.
  • Interleave A/B measurements (ABABAB, not AAABBB) so thermal drift hits both sides equally.

torch.utils.benchmark.Timer packages this discipline — synchronisation, warm-up, adaptive repetition, and blocked autorange — and is the reference implementation to imitate when rolling your own.

The event-pair pattern for pipelines

For a training step, bracket sub-phases with event pairs (data-copy, forward, backward, optimiser) and read all elapsed times once per N steps. Cost is a few microseconds per event; the resulting per-phase timeline catches regressions that end-to-end timing averages away, and overlaps cleanly with profiler traces when a deep dive is needed.

References

  • NVIDIA, CUDA C++ Programming Guide — streams, events, and the asynchronous execution contract.
  • NVIDIA, CUDA C++ Best Practices Guide — the timing methodology chapter.
  • Hoefler and Belli (2015), Scientific Benchmarking of Parallel Computing Systems, SC — the statistics your benchmark table should have used.

What to learn next