Skip to main content
Learn AI Create the Future — with Pranay
  • Learn
  • Paths
  • Projects
  • Practice
  • Glossary
  • Errors
  • LeetCode
  • Main site

Main site →
Select language

Translation is unavailable right now. Your browser’s own “Translate page” usually works.

Machine translation by Google

Topics

  • 🧱 Foundations 2 sections
  • 🤖 Machine Learning 11 sections
  • 📊 Statistics and Experiments 4 sections
  • 🧭 Doing the Work 9 sections
  • 🧠 Deep Learning 11 sections
    • Deep Learning 13 lessons
    • PyTorch Tensors 10 lessons
    • Autograd in Depth 10 lessons
    • Building Models with nn.Module 11 lessons
    • Datasets and DataLoaders 10 lessons
    • Optimisers, Schedulers and the Training Loop 9 lessons
    • Debugging PyTorch 9 lessons
    • GPU Memory and Speed 10 lessons
    • Multi-GPU and Distributed Training 11 lessons
      • Overview
      • DataParallel vs DistributedDataParallel
      • Your first DDP run with torchrun
      • DistributedSampler and set_epoch
      • When gradients are synchronised, and no_sync
      • Scaling the learning rate with the number of GPUs
      • Saving and logging from one rank only
      • all_reduce, all_gather and broadcast
      • FSDP: sharding a model across GPUs
      • Pipeline and tensor parallelism
      • Training across several machines
      • Debugging a distributed job that hangs
    • Checkpoints, Export and Inference 7 lessons
    • Reinforcement Learning 13 lessons
  • 🧰 Libraries and Frameworks 4 sections
  • 💬 Language and NLP 15 sections
  • 👁️ Computer Vision 13 sections
  • 🎙️ Speech, Forecasting and Recommenders 3 sections
  • ✨ Generative AI and LLMs 4 sections
  • 🔬 Inside a Transformer 11 sections
  • ⚙️ Production and MLOps 16 sections
  • 🌍 AI in the Real World 12 sections
  • ⚖️ Safety, Ethics and Law 2 sections

See all topics

  1. Home
  2. Learn
  3. Multi-GPU and Distributed Training

🧠 Deep Learning · Section 034

🛰️ Multi-GPU and Distributed Training

Training on more than one GPU: what actually gets sent between them, and the ways it goes wrong.

Every lesson in this section is written by Pranay Mahendrakar.

11 of 11 lessons published · Three reading levels on every lesson

Start with “DataParallel vs DistributedDataParallel”

Lessons in order

Work top to bottom. Each lesson assumes the one above it.

  1. 01 DataParallel vs DistributedDataParallel
  2. 02 Your first DDP run with torchrun
  3. 03 DistributedSampler and set_epoch
  4. 04 When gradients are synchronised, and no_sync
  5. 05 Scaling the learning rate with the number of GPUs
  6. 06 Saving and logging from one rank only
  7. 07 all_reduce, all_gather and broadcast
  8. 08 FSDP: sharding a model across GPUs
  9. 09 Pipeline and tensor parallelism
  10. 10 Training across several machines
  11. 11 Debugging a distributed job that hangs
Previous GPU Memory and Speed Next Checkpoints, Export and Inference

Learn AI

Learn Artificial Intelligence the easy way.

Learn. Build. Ask. Share.

Free forever. No sign-up, no ads, no tracking.

Learn

  • All topics
  • Learning paths
  • Projects
  • Practice playground
  • AI glossary
  • Error database
  • LeetCode solutions

Start here

  • Python for AI
  • Mathematics for AI
  • Machine Learning
  • Classic Algorithms in Depth
  • Linear Models and Regularisation
  • Ensembles and Gradient Boosting

This site

  • Search
  • New lessons feed
  • Sitemap
  • About Pranay Mahendrakar
  • pranaymahendrakar.com

© 2026 Pranay Mahendrakar. Written for people who are starting from zero.

10 new lessons every day. If something here is wrong or confusing, that is worth fixing — say so.