The HuggingFace Stack

Accelerate

Accelerate takes an ordinary PyTorch training loop and makes it run unchanged on a CPU, one GPU or eight, by replacing the device handling with prepare() and backward().

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Accelerate lets one training script run on whatever hardware you have, with no edits. A laptop CPU today, eight GPUs next month.

Think of a phone charger that works in India, Japan and Germany. The plug shape and the voltage differ everywhere, and the charger absorbs all of it. You carry one charger, not three, and you stop thinking about sockets.

Training code has the same problem. The instructions to run on a laptop, on one graphics card, or on a rack of them are annoyingly different. Accelerate is the universal adapter.

Why it exists

A GPU is the fast chip that does the heavy number work. Code that uses one is littered with instructions moving data onto it and results back off. Code for several GPUs adds a whole extra layer of coordination.

The result was three versions of every script, drifting apart, and a laptop version that nobody kept working. Changing hardware meant rewriting the loop.

Accelerate strips the hardware talk out of your loop and hides it behind two calls. The loop that remains looks the way training loops looked before GPUs existed.

How it works

your ordinary loop:
    for batch in data:
        loss = model(batch)
        loss.backward()
        step

with Accelerate:
    model, optimizer, data = accelerator.prepare(model, optimizer, data)
    for batch in data:                     ← batches arrive on the right device already
        loss = model(batch)
        accelerator.backward(loss)         ← the one changed line
        step

You never write "put this on the graphics card". prepare wraps your pieces, and every batch arrives where it should.

A real example you have seen

Google Docs. The same document opens on a phone, a laptop and a projector, and you never think about screen sizes while typing. The layout adapts; your writing does not change.

Remember this

  • One script, any hardware: Accelerate handles the differences.
  • prepare() wraps your model, optimizer and data loader.
  • accelerator.backward(loss) replaces loss.backward(). That is close to the whole change.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install accelerate torch

Tested with accelerate 1.13 and torch 2.5. No model download at all — the demos build a tiny network and random data, so everything runs on a laptop CPU in under a second.

A whole training loop with no device code in it

loop.py
import torch
from torch.utils.data import TensorDataset, DataLoader
from accelerate import Accelerator

torch.manual_seed(0)
X = torch.randn(256, 4)
y = ((X[:, 0] + X[:, 1]) > 0).long()           # a rule the model can actually find
loader = DataLoader(TensorDataset(X, y), batch_size=32, shuffle=True)

model = torch.nn.Sequential(torch.nn.Linear(4, 16), torch.nn.ReLU(), torch.nn.Linear(16, 2))
opt = torch.optim.AdamW(model.parameters(), lr=0.05)
loss_fn = torch.nn.CrossEntropyLoss()

acc = Accelerator()
model, opt, loader = acc.prepare(model, opt, loader)   # no .to(device) anywhere in this file
acc.print("device:", acc.device, "| processes:", acc.num_processes)

for epoch in range(3):
    running = 0.0
    for xb, yb in loader:
        loss = loss_fn(model(xb), yb)
        acc.backward(loss)                              # replaces loss.backward()
        opt.step(); opt.zero_grad()
        running += loss.item()
    acc.print(f"epoch {epoch}  loss {running / len(loader):.3f}")
Output
device: cpu | processes: 1
epoch 0  loss 0.479
epoch 1  loss 0.149
epoch 2  loss 0.077

Run the same file on a machine with a CUDA graphics card and the first line reads device: cuda. On the machine this was written on, the three loss numbers came out identical either way — the arithmetic did not change, only where it happened.

Counting correctly when there is more than one process

metrics.py
import torch
from torch.utils.data import TensorDataset, DataLoader
from accelerate import Accelerator

torch.manual_seed(0)
X = torch.randn(150, 4)
y = ((X[:, 0] + X[:, 1]) > 0).long()
loader = DataLoader(TensorDataset(X, y), batch_size=32)   # 150 rows, 32 each: the last batch is short

model = torch.nn.Sequential(torch.nn.Linear(4, 16), torch.nn.ReLU(), torch.nn.Linear(16, 2))
acc = Accelerator()
model, loader = acc.prepare(model, loader)

correct = seen = 0
model.eval()
with torch.no_grad():
    for xb, yb in loader:
        preds = model(xb).argmax(-1)
        preds, yb = acc.gather_for_metrics((preds, yb))   # correct on 1 process or on 8
        correct += (preds == yb).sum().item()
        seen += yb.numel()
acc.print(f"scored {seen} examples, accuracy {correct / seen:.2f}")
Output
scored 150 examples, accuracy 0.39

The model here was never trained, so 0.39 is coin-flip noise. The number that matters is 150. Not 152, not 160 — every example scored once.

The walkthrough

prepare returns replacements; it does not modify in place. Reassign every variable you pass in. Writing acc.prepare(model, opt, loader) without capturing the results leaves you training the unwrapped originals, on the wrong device, with no error message.

accelerator.backward(loss) exists for the invisible work. On plain CPU it calls loss.backward(). Under mixed precision it scales the loss first, and under gradient accumulation it divides by the accumulation count. Writing loss.backward() by hand works today and breaks silently the day you enable either — see backpropagation for what it is calling underneath.

gather_for_metrics is not the same as gather. Across several processes the data loader pads the last batch so every process gets an equal share. A plain gather counts those padding rows. gather_for_metrics trims them, which is why the count above is exactly the dataset size. Get this wrong and your accuracy is quietly computed over duplicated examples.

acc.print prints once, not once per process. On eight GPUs a plain print gives eight identical lines. acc.print prints only from the main process. The same idea covers acc.is_main_process for anything that should happen once — saving, uploading, writing a file.

Gradient accumulation is a context manager. Small GPU, want a large effective batch:

accumulate.py
import torch
from torch.utils.data import TensorDataset, DataLoader
from accelerate import Accelerator

torch.manual_seed(0)
X = torch.randn(64, 4); y = ((X[:, 0] + X[:, 1]) > 0).long()
loader = DataLoader(TensorDataset(X, y), batch_size=8)     # 64 rows / 8 = 8 batches
model = torch.nn.Linear(4, 2)
opt = torch.optim.SGD(model.parameters(), lr=0.1)

acc = Accelerator(gradient_accumulation_steps=4)
model, opt, loader = acc.prepare(model, opt, loader)

updates = 0
for xb, yb in loader:
    with acc.accumulate(model):
        loss = torch.nn.functional.cross_entropy(model(xb), yb)
        acc.backward(loss)
        if acc.sync_gradients:
            updates += 1
        opt.step(); opt.zero_grad()
acc.print("batches:", len(loader), "| optimizer updates:", updates)
Output
batches: 8 | optimizer updates: 2

opt.step() sits inside the loop, but it only updates on every fourth batch — eight batches, two updates. Accelerate also skips the cross-process gradient sync on the in-between batches, which is where most of the speed comes from.

Two commands run it on real hardware. accelerate config asks a short list of questions — how many machines, how many GPUs, which precision — and writes a config file. accelerate launch loop.py then starts the right number of processes with the right environment. Your script never changes. Inside a notebook, notebook_launcher(main, num_processes=2) does the same job.

Common mistakes

Calling .to("cuda") anywhere in the file. That hard-codes one device and defeats the point. If you truly need the device, it is acc.device — and after prepare, batches are already there.

Saving the wrapped model. Under distributed training model is wrapped in a DistributedDataParallel object, and its saved keys all gain a module. prefix that later refuses to load. Unwrap first: acc.unwrap_model(model).save_pretrained(path), guarded by acc.is_main_process.

Mixing up the two batch sizes. With 4 processes and batch_size=32, the effective batch is 128 and your learning rate is now wrong for it. Decide the effective batch first, then divide.

Forgetting acc.wait_for_everyone() before reading a file you wrote. Processes run at different speeds. One can reach the read before another has finished writing, which shows up as a corrupt-file error that never reproduces on one GPU.

Assuming it will make one GPU faster. Accelerate is portability, not speed. On a single machine, the throughput comes from mixed precision (Accelerator(mixed_precision="bf16")), a larger batch, or better data loading — see is the GPU waiting for data?.

Try it yourself

Add Accelerator(mixed_precision="bf16") to loop.py and rerun on CPU — read the error or warning you get and work out why. Then add gradient accumulation of 4 and print how many optimizer updates happen per epoch. Predict the number before you run it.

What to learn next

Researcher — Mathematics and papers.

What prepare actually returns

Each argument is wrapped by type:

  • Model: moved to acc.device, then wrapped by DistributedDataParallel, FSDP, or a DeepSpeed engine depending on config; forward is wrapped in an autocast context under mixed precision.
  • Optimizer: wrapped in AcceleratedOptimizer, which owns the gradient-scaler interaction and the accumulation-aware step skipping.
  • DataLoader: rebuilt around a BatchSamplerShard so process i sees a disjoint shard, plus a device-placement wrapper. The last batch is padded so every process performs the same number of steps — collectives deadlock otherwise, which is the underlying reason gather_for_metrics must exist to un-pad.
  • Scheduler: stepped num_processes times per optimizer step, so a schedule written for one process spans the same number of epochs rather than the same number of steps.

Accelerator reads its configuration from environment variables set by accelerate launch, so the same script yields DDP, FSDP, DeepSpeed, TPU or single-process execution with no source change. This is the same layer Trainer sits on, which is why TrainingArguments can expose FSDP and DeepSpeed settings without implementing either.

Data parallelism and its accounting

Under data parallelism each process holds a full model replica and a distinct shard of the batch; gradients are all-reduced before the step, so all replicas stay identical. The gradient computed is the mean over the global batch, making the effective batch per_device_batch × num_processes × accumulation_steps.

Two consequences are routinely missed. Learning rate should scale with effective batch — linear scaling with warmup (Goyal et al., 2017, Accurate, large minibatch SGD) is the standard starting rule, and the square-root rule is the common alternative. And batch normalisation statistics are computed per replica unless synchronised, so a small per-device batch changes the model's behaviour, not only its speed — see BatchNorm in PyTorch.

Gradient accumulation reproduces a large batch's gradient but not its wall-clock behaviour: memory is bounded by the micro-batch's activations, while step count falls by the accumulation factor. Accelerate's no_sync on non-boundary micro-batches removes the all-reduce that would otherwise be pure waste.

Where it stops being enough

Data parallelism requires one full replica per device — parameters, gradients and optimizer states, roughly 16 bytes per parameter for Adam in mixed precision before activations. Past that, the model must be sharded rather than replicated: ZeRO stages 1–3 (Rajbhandari et al., 2020) partition optimizer states, gradients and parameters in turn, and PyTorch FSDP implements the same decomposition natively. Tensor and pipeline parallelism split individual layers instead, and are configured through the same launcher.

Accelerate's own contribution at that scale is not the parallelism — it is the uniform surface. The loop you debugged on a laptop CPU is the loop that runs on the cluster, which is worth more than it sounds when a distributed run fails at 3 a.m.

What to learn next