Deep Learning Section 033
GPU Memory and Speed
Why training runs out of memory or crawls, and the small number of changes that reliably fix it.
10 of 10 lessons published Three reading levels on every lesson
Lessons in order
Work top to bottom. Each lesson assumes the one above it.
- Reading and fixing CUDA out of memory
- What is using your GPU memory
- Mixed precision with autocast and GradScaler
- Gradient checkpointing
- torch.compile
- Profiling with torch.profiler
- Timing GPU code correctly
- Fast attention with scaled_dot_product_attention
- channels_last for convolution networks
- When the GPU is slower than the CPU