Building Models with nn.Module

Padding and packing variable-length sequences

Sentences come in different lengths but tensors demand rectangles — padding fills the gaps with a placeholder, and packing tells the RNN to skip the filler entirely.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Padding makes unequal sentences the same length by adding filler; packing lets the model skip the filler instead of computing on it.

Picture a sleeper coach with rows of four seats, booked by families of different sizes. A family of two still occupies a four-seat row — two seats hold people, two hold luggage nobody cares about.

That is padding: empty seats filled with a placeholder so every row has the same shape.

Now the ticket collector walks through. A lazy collector checks all four seats in every row, wasting time on luggage. A smart one carries the passenger list and visits occupied seats only. That smart route is packing.

Why it exists

Models process sentences in groups — batches — because groups are fast. But a batch must be a rectangle: every row the same length. Language refuses. "Nice." is one word; some sentences run to forty.

So we pad: pick the longest sentence in the batch, and fill every shorter one with a placeholder token up to that length.

Padding creates a second problem. A model that reads left to right, like an LSTM, will happily read the filler too. Its final memory of a short sentence becomes a memory of trailing nothing. Packing fixes this: alongside the rectangle, we hand over each sentence's true length, and the model stops reading each row at the right place.

How it works

three sentences:      lengths:
  [A B C D E]            5
  [F G]                  2
  [H I J]                3

padded rectangle:          packed reading order:
  A B C D E                step 1: A F H   (3 rows still alive)
  F G . . .        →       step 2: B G I   (3 alive)
  H I J . .                step 3: C J     (2 alive — F G finished)
                           step 4: D       (1 alive)
                           step 5: E

The dots never get read. Nothing is computed on luggage.

A real example you have seen

Every voice assistant and translation app batches many utterances of different lengths at once. Padding and length-tracking are how one rectangular batch can hold your three words and someone's rambling paragraph.

Remember this

  • Batches must be rectangles; padding fills the short rows.
  • Packing carries the true lengths so recurrent models skip the filler.
  • Forgetting the lengths does not crash — it quietly teaches the model about luggage.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

Written and tested against torch 2.5 on CPU. Shapes and integer output below are exact.

Pad, pack, run, unpack

pack_pipeline.py
import torch
from torch import nn
from torch.nn.utils.rnn import pad_sequence, pack_padded_sequence, pad_packed_sequence

torch.manual_seed(0)
# Three sentences of different lengths, already turned into ids.
sentences = [torch.tensor([4, 2, 9, 7, 1]),
             torch.tensor([5, 3]),
             torch.tensor([8, 6, 2])]
lengths = torch.tensor([len(s) for s in sentences])

padded = pad_sequence(sentences, batch_first=True, padding_value=0)
print("padded batch:")
print(padded)

emb = nn.Embedding(10, 6, padding_idx=0)
lstm = nn.LSTM(input_size=6, hidden_size=8, batch_first=True)

packed = pack_padded_sequence(emb(padded), lengths, batch_first=True, enforce_sorted=False)
print("packed batch_sizes:", packed.batch_sizes.tolist())

out, (h, c) = lstm(packed)
out, out_lengths = pad_packed_sequence(out, batch_first=True)
print("output shape:", tuple(out.shape))
print("lengths came back:", out_lengths.tolist())
Output
padded batch:
tensor([[4, 2, 9, 7, 1],
        [5, 3, 0, 0, 0],
        [8, 6, 2, 0, 0]])
packed batch_sizes: [3, 3, 2, 1, 1]
output shape: (3, 5, 8)
lengths came back: [5, 2, 3]

The walkthrough

pad_sequence takes a plain list of unequal tensors and returns the rectangle, filling with padding_value=0 — matching the padding_idx of the embedding from the previous lesson. Keep those two numbers identical; that is the contract.

batch_first=True everywhere. PyTorch's historical default puts time first, (seq, batch, features), and mixing conventions across calls is the number one shape bug in sequence code. Pick batch-first and say it in every call.

batch_sizes: [3, 3, 2, 1, 1] is the smart ticket collector's plan. At time step 1, three sequences are alive; by step 3 the two-word sentence has finished, so 2 remain; steps 4 and 5 touch one row. Total work: 3+3+2+1+1 = 10 steps instead of 15 for the full rectangle.

enforce_sorted=False lets you pass sentences in any order — PyTorch sorts internally and remembers the permutation. The default True demands you pre-sort by descending length and exists for exporting models; day to day, pass False.

h versus out. After unpacking, out[i, -1] is the output at the padded last position — luggage for short rows. The final hidden state h is length-aware: h[-1, i] is sequence i's state at its true last step. For a sentence classifier, use h, or index out with the returned lengths.

Do transformers need this?

No — and it is worth knowing why. A transformer reads all positions at once, so there is nothing to stop early; instead it takes an attention mask telling it which positions are real. Padding survives; packing is specific to step-by-step models. The masking story lives in attention.

Common mistakes

Skipping packing entirely. Nothing crashes. The LSTM reads the zeros, and its final states for short sequences are polluted. Accuracy drops a little, mysteriously — this bug hides for months.

Padding value colliding with a real id. If 0 means "the" in your vocabulary, padding with 0 makes fake "the"s. Reserve id 0 for <pad> from the start.

Taking out[:, -1] as the final state. Correct only when every sequence in the batch has full length. Use h, as above.

Computing loss on padded positions. For per-token tasks, pass ignore_index to nn.CrossEntropyLoss with your pad id, or the model earns reward for predicting filler.

Try it yourself

Remove the packing — feed emb(padded) straight into the LSTM — and compare h for the two-word sentence against the packed version. The difference you see is exactly the pollution from reading three steps of luggage.

What to learn next

Researcher — Mathematics and papers.

What packing actually stores

A PackedSequence is two tensors: data, the concatenation of all alive rows in time-major order, and batch_sizes, the per-step count of alive sequences (plus permutation indices when unsorted input was allowed). It is a ragged-tensor encoding specialised for recurrence: step $t$ is a contiguous slice of data of height batch_sizes[t], so the RNN kernel runs one dense matmul per step with no wasted rows.

Total compute drops from $O(B \cdot T_{max})$ to $O(\sum_i T_i)$ — the padding-efficiency ratio is the mean-to-max length ratio of the batch. This is why length-bucketed batching (grouping similar-length sequences per batch) multiplies throughput for text: it raises that ratio toward 1. The cuDNN variable-length RNN kernels consume batch_sizes directly (the historical reason for the descending-length sort requirement).

The mask alternative

Transformers replaced packing with masking: attention logits at padded keys are set to $-\infty$ before softmax, giving zero weight. Compute is not saved — masked positions are computed then discarded — which at scale motivated a return of packing in a new costume: sequence packing for pretraining, concatenating multiple documents into one fixed-length row with block-diagonal attention masks (used in T5, and via attn_mask/varlen kernels in FlashAttention-2, Dao 2023). NestedTensor (torch.nested) is PyTorch's ongoing native ragged representation; as of torch 2.5 it remains in prototype/limited coverage — check current docs before building on it.

Gradient correctness

Packing is not only about speed. For an unpacked padded batch, the recurrent update at a padded step still propagates hidden state through time; even with zero inputs, biases and the recurrent weight matrix act, so $h_{T_{max}}$ for a short sequence differs from $h_{T_i}$ — and the gradients through those extra steps are pure artefact. Packing makes the artefact structurally impossible rather than approximately harmless. Empirically the damage of skipping this ranges from negligible (short pad tails) to several accuracy points (long-tailed length distributions) — an ablation worth actually running on your data.

Reading

  • Hochreiter and Schmidhuber (1997), Long Short-Term Memory — the model this machinery serves.
  • Dao (2023), FlashAttention-2 — varlen attention as packing's modern descendant.
  • PyTorch docs, torch.nn.utils.rnn — the authoritative statement of PackedSequence invariants.

What to learn next