Building Models with nn.Module

Sequential, ModuleList and ModuleDict

Three containers for holding layers — one runs them for you in a fixed line, and the other two register them while leaving the running to your own forward.

On this page 5
  1. Why they exist
  2. How to choose
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Sequential, ModuleList and ModuleDict are containers that hold layers and register all of them at once — they differ in who controls the order of work.

Picture a tiffin-packing line and a toolbox. On the packing line, every box moves through the same fixed stations: rice, curry, lid, seal. Nobody decides anything per box — the line is the recipe.

A toolbox is different. It holds your tools in one place, but you decide which tool to pick up, when, and how many times. And if the toolbox has labelled drawers, you can ask for a tool by name.

Sequential is the packing line. ModuleList is the toolbox. ModuleDict is the toolbox with labelled drawers.

Why they exist

The previous lesson showed the trap: layers stored in a plain Python list are invisible to PyTorch. They never train and never save.

But real models often need a collection of layers. A network might have twelve repeated blocks, or one output head per task. You need a container that PyTorch looks inside.

That is the whole job of these three. They are lists and dicts that register their contents — meaning every layer inside them lands in the model's official record.

How to choose

Is the data flow a straight line, same every time?
        |                       |
       yes                      no
        |                       |
   Sequential          Do you pick layers by position or by name?
                            |                |
                        position           name
                            |                |
                       ModuleList       ModuleDict

A real example you have seen

Translation apps handle many language pairs. A model behind one can keep a shared body plus one small head per language, stored by name — exactly the labelled-drawers pattern.

Remember this

  • All three containers register every layer inside them.
  • Sequential also runs the layers for you, in order. The other two leave forward to you.
  • A plain Python list of layers is the classic silent bug. These containers are the fix.

What to learn next

  • Buffers vs parameters — the second thing nn.Module registers, and when to use it.
  • Weight initialisation — what the numbers inside those registered layers start as.
  • CNN — the architecture where residual ModuleList loops became standard.

Developer — Code and libraries.

Setup

bash
pip install torch

Written and tested against torch 2.5 on CPU.

All three, side by side

containers.py
import torch
from torch import nn

torch.manual_seed(0)

# One straight pipeline: Sequential is the right container.
mlp = nn.Sequential(
    nn.Linear(4, 16),
    nn.ReLU(),
    nn.Linear(16, 1),
)

# Depth chosen by a number: ModuleList, because forward needs a loop.
class DeepNet(nn.Module):
    def __init__(self, depth):
        super().__init__()
        self.blocks = nn.ModuleList(nn.Linear(8, 8) for _ in range(depth))

    def forward(self, x):
        for block in self.blocks:
            x = torch.relu(block(x))
        return x

# Named heads: ModuleDict, looked up by string.
class MultiHead(nn.Module):
    def __init__(self):
        super().__init__()
        self.body = nn.Linear(4, 8)
        self.heads = nn.ModuleDict({
            "price": nn.Linear(8, 1),
            "ripeness": nn.Linear(8, 3),
        })

    def forward(self, x, task):
        return self.heads[task](torch.relu(self.body(x)))

def count(m):
    return sum(p.numel() for p in m.parameters())

print("Sequential params:", count(mlp))
print("DeepNet(3) params:", count(DeepNet(3)))
print("DeepNet(6) params:", count(DeepNet(6)))
x = torch.randn(2, 4)
print("price head out:", tuple(MultiHead()(x, "price").shape))
print("ripeness head out:", tuple(MultiHead()(x, "ripeness").shape))
Output
Sequential params: 97
DeepNet(3) params: 216
DeepNet(6) params: 432
price head out: (2, 1)
ripeness head out: (2, 3)

The differences that matter

Sequential has a built-in forward. Calling mlp(x) feeds x through each layer in order. You never write the loop. The cost of that convenience: the flow is fixed. No branches, no skip connections, no second input.

ModuleList has no forward at all. Calling one raises NotImplementedError. It registers layers and nothing more — the loop in your forward is the model logic. That loop can skip blocks, reuse blocks, or add the input back in (the residual pattern used by CNN backbones).

ModuleDict is ModuleList with names. Handy when a string chooses the path: one head per task, per language, per dataset.

Parameter names follow the container: blocks.0.weight, blocks.1.weight, heads.price.weight. Position or key becomes part of the saved name — renaming a dict key breaks old checkpoints.

The classic failure, one more time

python
class Bad(nn.Module):
    def __init__(self):
        super().__init__()
        self.blocks = [nn.Linear(8, 8) for _ in range(3)]   # plain list

torch.optim.SGD(Bad().parameters(), lr=0.1)
Output
ValueError: optimizer got an empty parameter list

This is the lucky version, where the list held every layer, so the model has zero registered parameters and the optimizer complains. If even one layer is registered elsewhere, there is no error — the listed layers stay random forever while the rest train.

Common mistakes

Using Sequential and then fighting it. If you find yourself slicing mlp[:2] and gluing outputs mid-pipeline, the flow is not a straight line. Switch to ModuleList and write the forward you actually want.

Indexing a ModuleList with a tensor. self.blocks[i] needs a plain Python int. Convert with int(i) if the index came out of a tensor.

Storing losses or metrics in a ModuleDict. Containers are for modules with parameters. A plain dict is right for everything else.

Believing Sequential is slower or faster. It is the same layers, called in the same order. Container choice is about code shape, not speed.

Try it yourself

Add a residual connection to DeepNet: change the loop body to x = x + torch.relu(block(x)). Confirm the parameter count does not change, and think about why.

What to learn next

  • Buffers vs parameters — the second thing nn.Module registers, and when to use it.
  • Weight initialisation — what the numbers inside those registered layers start as.
  • CNN — the architecture where residual ModuleList loops became standard.

Researcher — Mathematics and papers.

What a container adds, formally

All three subclass nn.Module and route their children into _modules, keyed by stringified index or dict key. That is the entire mechanism: registration by attribute path, as in the registration lesson. Sequential additionally defines forward(x) = f_n(...f_2(f_1(x))) — composition in registration order.

The expressivity gap is therefore exact: Sequential computes only chain compositions, while a hand-written forward over a ModuleList computes any function expressible in Python over the registered children. Residual networks (He et al., 2016, Deep Residual Learning for Image Recognition) are the canonical function outside Sequential's reach — h = x + f(x) needs access to x twice. DenseNets (Huang et al., 2017) widen the gap: every block consumes all previous outputs.

State dict stability

Container choice fixes the parameter namespace:

  • Sequential: positional keys — inserting a layer renumbers everything after it, invalidating checkpoints. nn.Sequential(OrderedDict([...])) with explicit names avoids this.
  • ModuleList: positional, same renumbering hazard.
  • ModuleDict: keys are stable under insertion, ordered by insertion (it preserves dict order).

For long-lived training runs, name stability is an engineering constraint on par with the architecture itself. Loading with load_state_dict(..., strict=False) papers over mismatches but silently drops tensors — audit its returned missing_keys and unexpected_keys.

Interaction with graph capture

torch.compile traces through container iteration without issue as of torch 2.5 — a Python loop over a ModuleList unrolls at trace time, so depth becomes a compile-time constant. Data-dependent container indexing (self.heads[task] where task varies per call) induces one graph per key: fine for a handful of heads, a recompilation storm for hundreds. Mixture-of-experts routing therefore avoids ModuleDict dispatch in favour of batched expert computation (Shazeer et al., 2017, Outrageously Large Neural Networks).

Reading

  • He et al. (2016), Deep Residual Learning for Image Recognition — why forward needed to be code, not a list.
  • Huang et al. (2017), Densely Connected Convolutional Networks.
  • Shazeer et al. (2017), Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

What to learn next

  • Buffers vs parameters — the second thing nn.Module registers, and when to use it.
  • Weight initialisation — what the numbers inside those registered layers start as.
  • CNN — the architecture where residual ModuleList loops became standard.