Accelerate
Accelerate takes an ordinary PyTorch training loop and makes it run unchanged on a CPU, one GPU or eight, by replacing the device handling with prepare() and backward().
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Accelerate lets one training script run on whatever hardware you have, with no edits. A laptop CPU today, eight GPUs next month.
Think of a phone charger that works in India, Japan and Germany. The plug shape and the voltage differ everywhere, and the charger absorbs all of it. You carry one charger, not three, and you stop thinking about sockets.
Training code has the same problem. The instructions to run on a laptop, on one graphics card, or on a rack of them are annoyingly different. Accelerate is the universal adapter.
Why it exists
A GPU is the fast chip that does the heavy number work. Code that uses one is littered with instructions moving data onto it and results back off. Code for several GPUs adds a whole extra layer of coordination.
The result was three versions of every script, drifting apart, and a laptop version that nobody kept working. Changing hardware meant rewriting the loop.
Accelerate strips the hardware talk out of your loop and hides it behind two calls. The loop that remains looks the way training loops looked before GPUs existed.
How it works
your ordinary loop:
for batch in data:
loss = model(batch)
loss.backward()
step
with Accelerate:
model, optimizer, data = accelerator.prepare(model, optimizer, data)
for batch in data: ← batches arrive on the right device already
loss = model(batch)
accelerator.backward(loss) ← the one changed line
stepYou never write "put this on the graphics card". prepare wraps your pieces, and every batch arrives where it should.
A real example you have seen
Google Docs. The same document opens on a phone, a laptop and a projector, and you never think about screen sizes while typing. The layout adapts; your writing does not change.
Remember this
- One script, any hardware: Accelerate handles the differences.
prepare()wraps your model, optimizer and data loader.accelerator.backward(loss)replacesloss.backward(). That is close to the whole change.
What to learn next
- Loading a model in 4-bit — the other way to fit training on the hardware you own.
- Moving tensors between CPU and GPU — the device handling Accelerate is hiding.
- Gradient accumulation — the mechanism, one level down.
Developer — Code and libraries.
Setup
pip install accelerate torchTested with accelerate 1.13 and torch 2.5. No model download at all — the demos build a tiny network and random data, so everything runs on a laptop CPU in under a second.
A whole training loop with no device code in it
import torch
from torch.utils.data import TensorDataset, DataLoader
from accelerate import Accelerator
torch.manual_seed(0)
X = torch.randn(256, 4)
y = ((X[:, 0] + X[:, 1]) > 0).long() # a rule the model can actually find
loader = DataLoader(TensorDataset(X, y), batch_size=32, shuffle=True)
model = torch.nn.Sequential(torch.nn.Linear(4, 16), torch.nn.ReLU(), torch.nn.Linear(16, 2))
opt = torch.optim.AdamW(model.parameters(), lr=0.05)
loss_fn = torch.nn.CrossEntropyLoss()
acc = Accelerator()
model, opt, loader = acc.prepare(model, opt, loader) # no .to(device) anywhere in this file
acc.print("device:", acc.device, "| processes:", acc.num_processes)
for epoch in range(3):
running = 0.0
for xb, yb in loader:
loss = loss_fn(model(xb), yb)
acc.backward(loss) # replaces loss.backward()
opt.step(); opt.zero_grad()
running += loss.item()
acc.print(f"epoch {epoch} loss {running / len(loader):.3f}")device: cpu | processes: 1 epoch 0 loss 0.479 epoch 1 loss 0.149 epoch 2 loss 0.077
Run the same file on a machine with a CUDA graphics card and the first line reads device: cuda. On the machine this was written on, the three loss numbers came out identical either way — the arithmetic did not change, only where it happened.
Counting correctly when there is more than one process
import torch
from torch.utils.data import TensorDataset, DataLoader
from accelerate import Accelerator
torch.manual_seed(0)
X = torch.randn(150, 4)
y = ((X[:, 0] + X[:, 1]) > 0).long()
loader = DataLoader(TensorDataset(X, y), batch_size=32) # 150 rows, 32 each: the last batch is short
model = torch.nn.Sequential(torch.nn.Linear(4, 16), torch.nn.ReLU(), torch.nn.Linear(16, 2))
acc = Accelerator()
model, loader = acc.prepare(model, loader)
correct = seen = 0
model.eval()
with torch.no_grad():
for xb, yb in loader:
preds = model(xb).argmax(-1)
preds, yb = acc.gather_for_metrics((preds, yb)) # correct on 1 process or on 8
correct += (preds == yb).sum().item()
seen += yb.numel()
acc.print(f"scored {seen} examples, accuracy {correct / seen:.2f}")scored 150 examples, accuracy 0.39
The model here was never trained, so 0.39 is coin-flip noise. The number that matters is 150. Not 152, not 160 — every example scored once.
The walkthrough
prepare returns replacements; it does not modify in place. Reassign every variable you pass in. Writing acc.prepare(model, opt, loader) without capturing the results leaves you training the unwrapped originals, on the wrong device, with no error message.
accelerator.backward(loss) exists for the invisible work. On plain CPU it calls loss.backward(). Under mixed precision it scales the loss first, and under gradient accumulation it divides by the accumulation count. Writing loss.backward() by hand works today and breaks silently the day you enable either — see backpropagation for what it is calling underneath.
gather_for_metrics is not the same as gather. Across several processes the data loader pads the last batch so every process gets an equal share. A plain gather counts those padding rows. gather_for_metrics trims them, which is why the count above is exactly the dataset size. Get this wrong and your accuracy is quietly computed over duplicated examples.
acc.print prints once, not once per process. On eight GPUs a plain print gives eight identical lines. acc.print prints only from the main process. The same idea covers acc.is_main_process for anything that should happen once — saving, uploading, writing a file.
Gradient accumulation is a context manager. Small GPU, want a large effective batch:
import torch
from torch.utils.data import TensorDataset, DataLoader
from accelerate import Accelerator
torch.manual_seed(0)
X = torch.randn(64, 4); y = ((X[:, 0] + X[:, 1]) > 0).long()
loader = DataLoader(TensorDataset(X, y), batch_size=8) # 64 rows / 8 = 8 batches
model = torch.nn.Linear(4, 2)
opt = torch.optim.SGD(model.parameters(), lr=0.1)
acc = Accelerator(gradient_accumulation_steps=4)
model, opt, loader = acc.prepare(model, opt, loader)
updates = 0
for xb, yb in loader:
with acc.accumulate(model):
loss = torch.nn.functional.cross_entropy(model(xb), yb)
acc.backward(loss)
if acc.sync_gradients:
updates += 1
opt.step(); opt.zero_grad()
acc.print("batches:", len(loader), "| optimizer updates:", updates)batches: 8 | optimizer updates: 2
opt.step() sits inside the loop, but it only updates on every fourth batch — eight batches, two updates. Accelerate also skips the cross-process gradient sync on the in-between batches, which is where most of the speed comes from.
Two commands run it on real hardware. accelerate config asks a short list of questions — how many machines, how many GPUs, which precision — and writes a config file. accelerate launch loop.py then starts the right number of processes with the right environment. Your script never changes. Inside a notebook, notebook_launcher(main, num_processes=2) does the same job.
Common mistakes
Calling .to("cuda") anywhere in the file. That hard-codes one device and defeats the point. If you truly need the device, it is acc.device — and after prepare, batches are already there.
Saving the wrapped model. Under distributed training model is wrapped in a DistributedDataParallel object, and its saved keys all gain a module. prefix that later refuses to load. Unwrap first: acc.unwrap_model(model).save_pretrained(path), guarded by acc.is_main_process.
Mixing up the two batch sizes. With 4 processes and batch_size=32, the effective batch is 128 and your learning rate is now wrong for it. Decide the effective batch first, then divide.
Forgetting acc.wait_for_everyone() before reading a file you wrote. Processes run at different speeds. One can reach the read before another has finished writing, which shows up as a corrupt-file error that never reproduces on one GPU.
Assuming it will make one GPU faster. Accelerate is portability, not speed. On a single machine, the throughput comes from mixed precision (Accelerator(mixed_precision="bf16")), a larger batch, or better data loading — see is the GPU waiting for data?.
Try it yourself
Add Accelerator(mixed_precision="bf16") to loop.py and rerun on CPU — read the error or warning you get and work out why. Then add gradient accumulation of 4 and print how many optimizer updates happen per epoch. Predict the number before you run it.
What to learn next
- Loading a model in 4-bit — the other way to fit training on the hardware you own.
- Moving tensors between CPU and GPU — the device handling Accelerate is hiding.
- Gradient accumulation — the mechanism, one level down.
Researcher — Mathematics and papers.
What prepare actually returns
Each argument is wrapped by type:
- Model: moved to
acc.device, then wrapped byDistributedDataParallel, FSDP, or a DeepSpeed engine depending on config; forward is wrapped in an autocast context under mixed precision. - Optimizer: wrapped in
AcceleratedOptimizer, which owns the gradient-scaler interaction and the accumulation-aware step skipping. - DataLoader: rebuilt around a
BatchSamplerShardso process i sees a disjoint shard, plus a device-placement wrapper. The last batch is padded so every process performs the same number of steps — collectives deadlock otherwise, which is the underlying reasongather_for_metricsmust exist to un-pad. - Scheduler: stepped
num_processestimes per optimizer step, so a schedule written for one process spans the same number of epochs rather than the same number of steps.
Accelerator reads its configuration from environment variables set by accelerate launch, so the same script yields DDP, FSDP, DeepSpeed, TPU or single-process execution with no source change. This is the same layer Trainer sits on, which is why TrainingArguments can expose FSDP and DeepSpeed settings without implementing either.
Data parallelism and its accounting
Under data parallelism each process holds a full model replica and a distinct shard of the batch; gradients are all-reduced before the step, so all replicas stay identical. The gradient computed is the mean over the global batch, making the effective batch per_device_batch × num_processes × accumulation_steps.
Two consequences are routinely missed. Learning rate should scale with effective batch — linear scaling with warmup (Goyal et al., 2017, Accurate, large minibatch SGD) is the standard starting rule, and the square-root rule is the common alternative. And batch normalisation statistics are computed per replica unless synchronised, so a small per-device batch changes the model's behaviour, not only its speed — see BatchNorm in PyTorch.
Gradient accumulation reproduces a large batch's gradient but not its wall-clock behaviour: memory is bounded by the micro-batch's activations, while step count falls by the accumulation factor. Accelerate's no_sync on non-boundary micro-batches removes the all-reduce that would otherwise be pure waste.
Where it stops being enough
Data parallelism requires one full replica per device — parameters, gradients and optimizer states, roughly 16 bytes per parameter for Adam in mixed precision before activations. Past that, the model must be sharded rather than replicated: ZeRO stages 1–3 (Rajbhandari et al., 2020) partition optimizer states, gradients and parameters in turn, and PyTorch FSDP implements the same decomposition natively. Tensor and pipeline parallelism split individual layers instead, and are configured through the same launcher.
Accelerate's own contribution at that scale is not the parallelism — it is the uniform surface. The loop you debugged on a laptop CPU is the loop that runs on the cluster, which is worth more than it sounds when a distributed run fails at 3 a.m.
What to learn next
- Loading a model in 4-bit — the other way to fit training on the hardware you own.
- Moving tensors between CPU and GPU — the device handling Accelerate is hiding.
- Gradient accumulation — the mechanism, one level down.