PyTorch basics
PyTorch is the library most neural networks are written in, and its central trick is keeping a record of every calculation so it can trace an error back to every weight that caused it.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
PyTorch is a free Python library for building neural networks, and its main trick is remembering every calculation it did.
Think about checking a long restaurant bill. The total looks too high, so you run your finger down the itemised list. Two extra teas, one dish you never ordered. You found the cause because every line was written down.
PyTorch keeps that itemised list for arithmetic. Every multiplication, every addition, in order.
When the model's answer is wrong, PyTorch walks back down that list. It works out how much each number contributed to the mistake.
Why this had to exist
A model learns by adjusting its weights, and to adjust a weight you need to know how it affected the error. That calculation is backpropagation.
Doing it by hand is possible for a small network. People did it for years, and it is worth doing once.
It stops being possible at scale. A modern model has billions of weights and dozens of layer types. Nobody is deriving those corrections on paper, and every hand-derived formula is a bug waiting to happen.
PyTorch does it for you, correctly, every time. That single feature is called autograd — automatic gradient computation.
How it works
your numbers → a calculation → a wrong answer
↑ |
| ↓
| "how wrong?" (the loss)
| |
└──── gradients ← walk the list backwards ←┘
then: nudge every number a little, and repeatYou write the forward direction, the part that produces an answer. PyTorch builds the backward direction for you from the record it kept.
The five things you will actually use
- Tensor — a grid of numbers, like a NumPy array, that can also live on a graphics card.
- Autograd — the record-keeping that produces corrections.
- Module — a reusable block, such as one layer, that holds its own weights.
- Loss — the number saying how wrong the answer was. See loss functions.
- Optimizer — the thing that applies the corrections.
Everything else in PyTorch is built out of those five.
Where you have already seen it
Most of the AI you have touched was built with PyTorch or something close to it. Image generators, speech-to-text on your phone, translation, chat assistants, medical imaging tools.
It began at Meta and is now run by a shared foundation, so no single company controls it. It is free and open, and you can read the source of any part you doubt.
What is honestly hard here
PyTorch is not the hard part of machine learning.
You can learn its core in a weekend. Getting good data, deciding what to train, and knowing whether the result is any good will take you much longer. The library will run a bad idea perfectly.
Also plan for the download. Installing PyTorch pulls a few hundred megabytes. On a metered mobile connection, do it once, on wifi, before you need it.
Remember this
- PyTorch records every calculation so it can compute corrections automatically.
- The five pieces are tensor, autograd, module, loss and optimizer.
- The library is the easy part. The data and the judgement are the hard parts.
What to learn next
- Backpropagation — the algorithm autograd is automating for you.
- Loss functions — choosing the number your training loop pushes down.
- Hugging Face — pretrained models you can load into PyTorch in three lines.
Developer — Code and libraries.
Setup
pip install torchEverything below runs on CPU in seconds. If you have no graphics card, nothing here changes.
For a smaller download on a machine with no GPU, install the CPU-only build from https://download.pytorch.org/whl/cpu. It skips the CUDA libraries, which are most of the size.
Tensors
A tensor is a NumPy array that also knows how to compute its own gradients and how to move to a GPU.
import torch
a = torch.tensor([[1.0, 2.0], [3.0, 4.0]])
b = torch.ones(2, 2)
print("a shape:", tuple(a.shape), " dtype:", a.dtype)
print("a + b:\n", a + b)
print("a @ b:\n", a @ b)
print("a.mean():", a.mean().item())
print("device:", a.device)a shape: (2, 2) dtype: torch.float32
a + b:
tensor([[2., 3.],
[4., 5.]])
a @ b:
tensor([[3., 3.],
[7., 7.]])
a.mean(): 2.5
device: cpuTwo details that bite later. The default dtype is float32, not float64 as in NumPy — mixing the two raises a dtype error. And .item() pulls a single-element tensor out as a plain Python number, which you need whenever you print or store a value.
Autograd, in six lines
import torch
x = torch.tensor(3.0, requires_grad=True) # requires_grad tells PyTorch to keep the receipt
y = x ** 2 + 4 * x
print("y :", y.item())
y.backward() # walk the receipt backwards
print("dy/dx at x=3 :", x.grad.item())
print("x itself :", x.item())y : 21.0 dy/dx at x=3 : 10.0 x itself : 3.0
Check it by hand. The slope of x² + 4x is 2x + 4, which is 10 at x = 3. PyTorch derived that from the operations you wrote, not from a formula you supplied.
x.grad holds the answer. x itself is untouched — computing a gradient never changes a value. Moving the value is the optimizer's job.
A complete training loop
This is the shape of every PyTorch program you will ever write.
import torch
import torch.nn as nn
torch.manual_seed(0) # same starting weights every run
# Eight rooms. Input: size in hundreds of square feet. Target: rent in thousands.
X = torch.tensor([[1.0], [1.5], [2.0], [2.5], [3.0], [3.5], [4.0], [4.5]])
y = torch.tensor([[5.0], [6.5], [8.0], [9.5], [11.0], [12.5], [14.0], [15.5]])
model = nn.Linear(1, 1) # one input, one output: weight and bias
lossfn = nn.MSELoss()
opt = torch.optim.SGD(model.parameters(), lr=0.05)
for epoch in range(1, 401):
pred = model(X)
loss = lossfn(pred, y)
opt.zero_grad() # clear last step's gradients, they add up otherwise
loss.backward() # fill every parameter's .grad
opt.step() # move each parameter one step downhill
if epoch in (1, 100, 200, 400):
print(f"epoch {epoch:3d} loss {loss.item():8.4f}")
w = model.weight.item()
b = model.bias.item()
print()
print(f"learned rule: rent = {w:.3f} x size + {b:.3f}")
print(f"true rule : rent = 3.000 x size + 2.000")epoch 1 loss 106.6251 epoch 100 loss 0.0022 epoch 200 loss 0.0001 epoch 400 loss 0.0000 learned rule: rent = 3.001 x size + 1.998 true rule : rent = 3.000 x size + 2.000
The loss values may vary in the last digits on a different PyTorch release, because the random initialisation differs. The learned rule lands in the same place regardless.
Memorise the three lines in the middle, in that order:
opt.zero_grad() → loss.backward() → opt.step()Clear, compute, move. Every training script in every repository you read has this inside it somewhere.
Why zero_grad is not optional
PyTorch adds new gradients to whatever is already in .grad. That is deliberate, and it is how you split a large batch across several passes. It is also the most common silent bug in beginner code.
import torch
x = torch.tensor(2.0, requires_grad=True)
for step in range(1, 4):
y = x ** 2 # the true slope at x=2 is 4, every single time
y.backward()
print(f"step {step}: x.grad = {x.grad.item()}")
print()
x.grad = None # what opt.zero_grad() does for you
y = x ** 2
y.backward()
print("after clearing:", x.grad.item())step 1: x.grad = 4.0 step 2: x.grad = 8.0 step 3: x.grad = 12.0 after clearing: 4.0
The slope never changed. The stored number grew because each backward() added to it. Forget zero_grad() in a training loop and your effective step size climbs every epoch until the loss explodes.
Building your own module
nn.Linear is fine for a straight line. Anything real needs a class.
import torch
import torch.nn as nn
class SmallNet(nn.Module):
def __init__(self, in_features, hidden, out_features):
super().__init__() # required, or parameters are not registered
self.fc1 = nn.Linear(in_features, hidden)
self.fc2 = nn.Linear(hidden, out_features)
def forward(self, x):
return self.fc2(torch.relu(self.fc1(x)))
net = SmallNet(4, 16, 2)
print(net)
print("parameters:", sum(p.numel() for p in net.parameters()))
print("output shape:", tuple(net(torch.randn(5, 4)).shape))SmallNet( (fc1): Linear(in_features=4, out_features=16, bias=True) (fc2): Linear(in_features=16, out_features=2, bias=True) ) parameters: 114 output shape: (5, 2)
Assigning nn.Linear to self.fc1 registers its weights automatically, which is why net.parameters() finds them. Put layers in a plain Python list instead and they vanish from parameters(), so the optimizer never trains them and no error is raised. Use nn.ModuleList or nn.Sequential for that.
Call the model as net(x), never net.forward(x). The first runs registered hooks; the second skips them, which breaks quantisation and profiling tools later.
Common mistakes
Forgetting opt.zero_grad(). Shown above. The loss creeps upward for no visible reason.
Calling .backward() twice on one graph. You get RuntimeError: Trying to backward through the graph a second time. The graph is freed after the first backward pass. If you genuinely need two, pass retain_graph=True — but first check you did not mean to recompute the forward pass.
Accumulating the loss tensor instead of a float. total += loss keeps the whole computation graph alive for every batch in the epoch, and memory climbs until the process is killed. Write total += loss.item().
Forgetting model.eval() and torch.no_grad() at evaluation time. eval() switches dropout off and makes batch-norm use its running statistics. no_grad() stops the graph being built, which roughly halves memory and speeds things up. Missing eval() gives you quietly worse validation numbers.
Mixing devices. Expected all tensors to be on the same device means your data is on CPU and your model on GPU, or the reverse. Move both: model.to(device) and x = x.to(device).
Shape errors from a missing batch dimension. A single example of 4 features has shape (4,), and nn.Linear wants (batch, 4). Use x.unsqueeze(0) to add the batch dimension.
Try it yourself
In train.py, raise lr to 0.5. The loss becomes nan and the learned rule becomes nan with it — the steps were large enough to throw the weights to infinity. Then drop lr to 0.001: it survives, but after 400 epochs it is still short of the true rule. Finding the middle is real work, not a formality.
Then replace nn.Linear(1, 1) with the SmallNet above, sized SmallNet(1, 16, 1). It has 114 parameters instead of 2 for a problem that needs 2. Compare the final loss. Watching a bigger model do no better on a simple problem is a useful thing to have seen.
What to learn next
- Backpropagation — the algorithm autograd is automating for you.
- Loss functions — choosing the number your training loop pushes down.
- Hugging Face — pretrained models you can load into PyTorch in three lines.
Researcher — Mathematics and papers.
What autograd actually computes
PyTorch implements reverse-mode automatic differentiation. It is not symbolic differentiation and it is not numerical differencing; it applies the chain rule to a recorded sequence of primitive operations, each of which ships a hand-written derivative rule.
For f: R^n → R^m decomposed as f = f_L ∘ ... ∘ f_1, the Jacobian factorises as J = J_L · ... · J_1. Reverse mode evaluates this right-to-left as a sequence of vector–Jacobian products:
backward(v) computes vᵀ J = ((vᵀ J_L) J_{L−1}) ... J_1v is the incoming cotangent — the gradient argument to .backward(), defaulting to 1.0 for a scalar loss. No full Jacobian is ever materialised. That is why .backward() on a non-scalar raises grad can be implicitly created only for scalar outputs: with m > 1 there is no canonical v.
Cost asymmetry is the reason reverse mode is used for training. Forward mode costs O(n) passes for n inputs; reverse mode costs O(m) passes for m outputs. Training has one scalar output and billions of inputs, so reverse mode wins by a factor of the parameter count.
Cost model
Let F be the FLOP count of the forward pass.
forward ≈ F
backward ≈ 2F (grad wrt inputs, plus grad wrt parameters)
training step ≈ 3F (hence the "6 FLOPs per parameter per token" rule)
activation memory ≈ O(depth × batch × width)Activation memory, not parameter memory, is usually what runs out first. Gradient checkpointing (Chen et al., 2016) stores only a subset of activations and recomputes the rest during backward, trading roughly 30% extra compute for O(√L) memory in L layers. In PyTorch: torch.utils.checkpoint.checkpoint.
The tape is dynamic
The graph is built during execution and discarded after .backward(). This is define-by-run, and it is the design choice that separated PyTorch from the static-graph frameworks of the same era. Control flow is ordinary Python: a loop with a data-dependent bound records a different graph on each call, at no cost in expressiveness.
The price is that nothing is known ahead of time, which forecloses whole-graph optimisation. torch.compile (PyTorch 2.0, 2023) recovers it: TorchDynamo captures Python bytecode into an FX graph, guarding on the assumptions it made; AOTAutograd traces the backward pass; TorchInductor lowers both to Triton kernels on GPU or C++/OpenMP on CPU. Fusion of elementwise chains is where most of the gain comes from. Graph breaks — triggered by unsupported constructs or by a .item() call forcing a device sync — silently return you to eager execution; torch._dynamo.explain locates them.
Numerics worth knowing
Mixed precision. bfloat16 keeps float32's 8 exponent bits and cuts the mantissa to 7, so it holds the same dynamic range and rarely needs loss scaling. float16 has a narrower range and does, which is what torch.cuda.amp.GradScaler provides. Master weights and optimizer state are conventionally kept in float32.
Optimizer memory. Adam stores two moments per parameter. In float32 that is 4 bytes for the parameter, 4 for the gradient and 8 for the moments, so roughly 16 bytes per parameter before activations. A 7B-parameter model therefore needs about 112 GB for a naive full fine-tune, which is the arithmetic that motivates LoRA.
Determinism. torch.manual_seed fixes initialisation and sampling. It does not fix reduction order in multi-threaded or GPU kernels. Full determinism requires torch.use_deterministic_algorithms(True), CUBLAS_WORKSPACE_CONFIG=:4096:8, and accepting that some kernels have no deterministic implementation and will raise instead.
Distributed training
DDP replicates the model on every device and all-reduces gradients after backward, overlapping communication with computation by bucketing. Communication volume per step is the parameter count, independent of world size, using ring all-reduce.
FSDP shards parameters, gradients and optimizer state across devices, all-gathering each layer's parameters immediately before use and freeing them after. This is the ZeRO-3 partitioning of Rajbhandari et al. (2020). Memory per device falls roughly by the world size; communication rises, so it earns its cost only when the model does not fit under DDP.
Papers and references
- Paszke et al., PyTorch: An Imperative Style, High-Performance Deep Learning Library, NeurIPS 2019 — arxiv.org/abs/1912.01703
- Ansel et al., PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation, ASPLOS 2024.
- Chen et al., Training Deep Nets with Sublinear Memory Cost, 2016 — arxiv.org/abs/1604.06174
- Rajbhandari et al., ZeRO: Memory Optimizations Toward Training Trillion Parameter Models, 2019 — arxiv.org/abs/1910.02054
- Griewank and Walther, Evaluating Derivatives, SIAM, 2nd edition, 2008 — the standard reference on automatic differentiation.
- Baydin et al., Automatic Differentiation in Machine Learning: a Survey, JMLR 2018 — arxiv.org/abs/1502.05767
What to learn next
- Backpropagation — the algorithm autograd is automating for you.
- Loss functions — choosing the number your training loop pushes down.
- Hugging Face — pretrained models you can load into PyTorch in three lines.