Deep Learning Section 034
Multi-GPU and Distributed Training
Training on more than one GPU: what actually gets sent between them, and the ways it goes wrong.
11 of 11 lessons published Three reading levels on every lesson
Lessons in order
Work top to bottom. Each lesson assumes the one above it.
- DataParallel vs DistributedDataParallel
- Your first DDP run with torchrun
- DistributedSampler and set_epoch
- When gradients are synchronised, and no_sync
- Scaling the learning rate with the number of GPUs
- Saving and logging from one rank only
- all_reduce, all_gather and broadcast
- FSDP: sharding a model across GPUs
- Pipeline and tensor parallelism
- Training across several machines
- Debugging a distributed job that hangs