Building Models with nn.Module

nn.Embedding and index errors

nn.Embedding is a lookup table that turns id numbers into learned vectors — and "index out of range in self" means you asked for a row the table does not have.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

nn.Embedding is a table that swaps an id number for a list of learned numbers describing that thing.

Think of a cloakroom at a wedding hall. You hand over a token — number 27 — and the attendant walks to shelf 27 and hands you what is stored there. The token itself means nothing. It is only a shelf position.

An embedding layer is that cloakroom. Words, users, or products are first given id numbers. The layer keeps one shelf per id, and each shelf holds that item's description — a vector, a short list of numbers the model refines during training.

And the famous error is the cloakroom's nightmare: someone hands over token 57, and there are only 50 shelves.

Why it exists

Models compute with numbers that have size — bigger means more. Raw ids are not like that. Word number 500 is not "more" than word number 3; the numbering is arbitrary, like phone numbers.

Feeding raw ids into a network would teach it patterns in the numbering, which are meaningless. What we want instead: each item gets its own learned description, and similar items end up with similar descriptions. During training, "tea" and "chai" drift close together; "tea" and "gearbox" drift apart.

That idea — meaning as position in number-space — is the embeddings story. nn.Embedding is its PyTorch form: a table of descriptions, looked up by id, improved by training.

How it works

"the mango is ripe"
        |
   [ tokenizer assigns ids ]        →  1, 2, 3, 4
        |
   [ embedding table, 5 rows ]
     row 1 → [ 0.2, -1.1,  0.7,  0.3 ]
     row 2 → [ 1.4,  0.2, -0.5,  0.9 ]   ...one vector per id
        |
   vectors flow into the rest of the model

Ask for row 5 in a 5-row table (rows 0 to 4) and the lookup fails. That is the whole anatomy of the error.

A real example you have seen

Every chatbot begins exactly this way: your words become ids, ids become vectors, vectors flow onward. Recommendation systems do it with people and products — your user id has a shelf in a table somewhere.

Remember this

  • An embedding is a lookup table: id in, learned vector out.
  • Ids are arbitrary labels; the learned vectors carry the meaning.
  • "Index out of range" means an id was too big for the table. Count your ids honestly.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

Written and tested against torch 2.5 on CPU.

The table, the lookup, and the crash

embedding_basics.py
import torch
from torch import nn

torch.manual_seed(0)
vocab = {"<pad>": 0, "the": 1, "mango": 2, "is": 3, "ripe": 4}
emb = nn.Embedding(num_embeddings=len(vocab), embedding_dim=4, padding_idx=0)

ids = torch.tensor([[1, 2, 3, 4],
                    [1, 2, 0, 0]])          # second sentence padded to length 4
out = emb(ids)
print("input ids shape: ", tuple(ids.shape))
print("output shape:    ", tuple(out.shape))
print("<pad> row:", out[1, 2].tolist())
try:
    emb(torch.tensor([5]))                  # one past the end of a size-5 table
except IndexError as e:
    print("IndexError:", e)
Output
input ids shape:  (2, 4)
output shape:     (2, 4, 4)
<pad> row: [0.0, 0.0, 0.0, 0.0]
IndexError: index out of range in self

The walkthrough

num_embeddings is the row count — how many distinct ids the table serves. embedding_dim is the shelf width — how many numbers describe each item. The table is one weight matrix of shape (num_embeddings, embedding_dim), fully trainable, registered like any parameter from the registration lesson.

The shape rule: output shape is input shape plus one final dimension of size embedding_dim. Ids (2, 4) became vectors (2, 4, 4). No flattening, no batching rules to memorise.

padding_idx=0 declares row 0 as the filler row for short sentences. That row is fixed at zeros and receives no gradient — filler should mean nothing and learn nothing. The next lesson builds on this.

Ids must be integer tensors. Passing floats raises a dtype error. Tokenizers hand you integers already; trouble usually starts when ids pass through NumPy or pandas and come back as floats.

Decoding "index out of range in self"

The message names neither the bad id nor the table size, so find them yourself:

python
print(ids.min().item(), ids.max().item(), emb.num_embeddings)

The rule the data must satisfy: every id in 0 <= id <= num_embeddings - 1. The three classic ways to break it:

  1. Off-by-one at the top. Vocabulary has ids 0 to 4999, table built with num_embeddings=4999. The largest legal row is 4998. Build with the count (5000), not the largest id.
  2. Special tokens counted late. You add <pad>, <unk>, <eos> after sizing the table. Three ids too many.
  3. A different tokenizer at inference. The model shipped with one vocabulary; the serving code tokenizes with another. Ids beyond the table arrive only on rare words, so it crashes occasionally in production — the worst kind.

On CPU the crash is a clean IndexError. On GPU it may surface as CUDA error: device-side assert triggered, reported at some later line because GPU work is asynchronous. Same bug, uglier costume — rerun on CPU to see it plainly.

Common mistakes

Sizing the table from data you have, not data you will see. Build from the full vocabulary and map unknown words to <unk>, or unseen test words will crash the lookup.

Forgetting padding_idx. Training still works, but the filler token gets a trained vector, and the model learns things about the padding. Wasteful and subtly wrong.

Confusing the two dimensions. nn.Embedding(4, 50000) builds 4 rows of width 50000. It crashes on id 4 immediately. The signature is rows first, width second.

One-hot encoding then a Linear layer. Mathematically the same operation, vastly more memory. The embedding lookup is that multiplication, done without building the one-hot matrix.

Try it yourself

Add "sweet": 5 to the vocabulary but leave num_embeddings at 5, and run a sentence containing it. Then fix the size and confirm. You will recognise this crash instantly for the rest of your career.

What to learn next

Researcher — Mathematics and papers.

The lookup as sparse matrix multiplication

For vocabulary size $V$ and dimension $d$, the table is $E \in \mathbb{R}^{V \times d}$. A lookup of id $i$ equals $e_i^\top E$, where $e_i$ is the one-hot basis vector — so nn.Embedding is nn.Linear (bias-free) applied to one-hot inputs, specialised to skip materialising them. The backward pass is correspondingly sparse: a batch touching $k$ distinct ids produces gradient rows for exactly those $k$ rows. nn.Embedding(sparse=True) returns a genuinely sparse gradient, which pairs only with sparse-aware optimizers (SparseAdam, SGD); dense Adam on a sparse gradient errors out.

The memory arithmetic dominates real systems: a 50,000-word table at $d = 768$ in fp32 is 153 MB; recommendation-scale tables (billions of ids) exceed single-node memory entirely, motivating hashed embeddings, quotient-remainder tricks (Shi et al., 2020), and embedding sharding — in industrial recommenders the embedding tables, not the network, hold most parameters.

padding_idx semantics

The designated row is zero-initialised and its gradient is masked — implemented as a hard zeroing of that row's gradient, not removal from the graph. A manual write (emb.weight.data[0] = ...) sticks, since only gradients are blocked. Downstream layers still see the zero vectors: masking attention or packing sequences (next lesson) is a separate concern from freezing the pad row.

Initialisation and scale

Default init is $\mathcal{N}(0, 1)$ per element — notably not fan-in scaled, unlike nn.Linear (see weight initialisation). The original Transformer (Vaswani et al., 2017) multiplied embedding outputs by $\sqrt{d}$ to balance against positional encodings; modern implementations vary, and the constant interacts with learning-rate choice. For tied input/output embeddings — sharing $E$ with the pre-softmax projection (Press and Wolf, 2017, Using the Output Embedding to Improve Language Models) — the scale question doubles: tying saves $Vd$ parameters and consistently improves perplexity at small scale.

Historical line

Learned distributed word representations enter with Bengio et al. (2003), A Neural Probabilistic Language Model; word2vec (Mikolov et al., 2013) made pre-trained tables ubiquitous; contextual models (Peters et al., 2018, ELMo; Devlin et al., 2019, BERT) demoted the static table to layer zero of something deeper. The nn.Embedding mechanics are identical throughout — what changed is everything stacked on top.

Reading

  • Bengio et al. (2003); Mikolov et al. (2013) — the lineage.
  • Press and Wolf (2017) — weight tying.
  • Shi et al. (2020), Compositional Embeddings Using Complementary Partitions — tables past memory limits.

What to learn next