Debugging PyTorch

Finding label and target bugs

When the model and the code are both fine, the labels can still be wrong — off by one, misaligned with their features, or scrambled — and three small audits expose all of it.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have seen this
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A label bug means the answers your model learns from are wrong — even though every answer sheet, on its own, looks perfectly fine.

Picture a teacher grading forty exam papers against an answer key. But the key's pages got shuffled, so paper 7 is graded with paper 12's answers. Every paper is genuine. The key is genuine. The pairing is destroyed, and every grade is nonsense.

No error message can catch this. The papers are the right shape, the key is the right shape, and the grading process runs beautifully. Only the outcome — grades that make no sense — hints that something upstream broke.

Why it exists

Labels travel a long, fragile road before training: exported from tools, joined across files, sorted, filtered, split, converted between numbering systems. Each hop is a chance to shift by one, reorder one side, or renumber classes differently than the model expects. Data labelling covers how labels are born; this lesson is about how they die in transit.

And a model trained on scrambled answers does not fail loudly. It trains, converges — to mediocrity — and wastes weeks pointing suspicion at architectures and learning rates.

How it works

audit 1: are the label VALUES legal?    (right range, sensible counts)
audit 2: are labels PAIRED with their   (break the pairing on purpose:
         features?                       does performance collapse?)
audit 3: do a few samples LOOK right?   (open 10 and check by eye)

Three audits, none longer than a few lines, run before long training.

Where you have seen this

Every mixed-up delivery — the neighbour's parcel with your name sticker — is a pairing bug. Warehouses fight it with barcode checks at every hop, not with faith. Treat your labels with warehouse suspicion.

Remember this

  • Label bugs produce no error message — only quiet mediocrity.
  • The three audits: legal values, intact pairing, eyeball ten samples.
  • Run them before training, not after disappointment.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

Captured with torch 2.5.1 on CPU; seeded and deterministic.

The one label bug that IS loud

Class labels numbered 1..N — from a spreadsheet, an R export, a label tool that counts from one — meeting a model that counts from zero:

off_by_one.py
import torch
import torch.nn as nn

logits = torch.randn(4, 3)                   # a 3-class model: valid labels are 0, 1, 2
labels = torch.tensor([1, 3, 2, 1])          # someone's labels run 1..3, not 0..2
loss = nn.CrossEntropyLoss()(logits, labels)
Output
IndexError: Target 3 is out of bounds.

Merciful, when it fires. The cruel variant: labels 1..3 with num_classes=4 — every label legal, class 0 unused, no error ever, and the model quietly learns a shifted world. On CUDA this error also wears a scarier costume (device-side assert triggered, reported on some unrelated later line); rerun on CPU for the honest message, a trick from the device lesson.

label_audit.py
import torch

labels = torch.tensor([1, 3, 2, 1, 2, 3, 3, 1])      # suspect labels, straight from a CSV
num_classes = 3

print("min:", labels.min().item(), " max:", labels.max().item())
print("counts per value:", torch.bincount(labels).tolist())
assert labels.min() >= 0 and labels.max() < num_classes, \
    f"labels must be 0..{num_classes - 1}, found {labels.min()}..{labels.max()}"
Output
min: 1  max: 3
counts per value: [0, 3, 2, 3]
AssertionError: labels must be 0..2, found 1..3

Three lines of printing expose the range, the per-class counts (a zero count for class 0 is the off-by-one fingerprint), and the assert turns tomorrow's silent version of this bug into today's loud one. Leave it in the pipeline permanently.

Audit 2: is the pairing intact?

The subtle killer — features and labels each fine, the alignment between them broken (one side shuffled, sorted, or filtered without the other). The test: destroy the pairing on purpose and compare.

pairing_test.py
import torch
import torch.nn as nn

torch.manual_seed(0)
X = torch.randn(400, 12)
y_real = (X[:, 0] + X[:, 5] > 0).long()
y_broken = y_real[torch.randperm(len(y_real))]   # deliberately destroy the pairing

def quick_fit(X, y, steps=300):
    torch.manual_seed(1)                         # same start for both runs
    model = nn.Linear(12, 2)
    opt = torch.optim.Adam(model.parameters(), lr=0.05)
    for _ in range(steps):
        opt.zero_grad()
        nn.CrossEntropyLoss()(model(X), y).backward()
        opt.step()
    return (model(X).argmax(dim=1) == y).float().mean().item()

print(f"accuracy, real pairing:      {quick_fit(X, y_real):.1%}")
print(f"accuracy, shuffled pairing:  {quick_fit(X, y_broken):.1%}")
Output
accuracy, real pairing:      99.3%
accuracy, shuffled pairing:  60.5%

A big gap means the true pairing carries real signal — healthy. The alarming outcome is the other one: if your real labels score no better than deliberately scrambled ones, your "real" pairing is as good as random, and the bug lives wherever features and labels were last joined. (Scrambled still beats 50% here because a 400-sample memoriser picks up crumbs — that gap shrinks on held-out data.)

Common mistakes

Sorting one side. sorted(image_paths) next to labels read in CSV order — alignment destroyed at birth. Keep features and labels in one structure (one dataframe, one list of pairs) so they cannot be reordered separately.

Two label mappings. Train builds {"cat": 0, "dog": 1} from its own folder scan; validation's scan meets an extra folder and maps differently. Build the class-to-index mapping once, save it, load it everywhere.

Regression targets in mixed units. Half the rows in lakhs, half in rupees; or unnormalised targets producing huge losses that read like a learning-rate problem. Print y.min(), y.max(), y.mean() — the audit-1 habit works for regression too.

Trusting accuracy on the training set. Memorisation hides label noise: the one-batch test passes with completely random labels. Only held-out performance, and audits like these, see label problems.

Try it yourself

Take audit 2 and evaluate both models on 200 fresh samples generated by the same rule, instead of the training X. Predict both numbers first. The shuffled model should now land near 50% — write one sentence on why the crumbs vanished.

What to learn next

Researcher — Mathematics and papers.

Chance level as a statistical reference

The pairing test is an informal permutation test: shuffling $y$ samples from the null distribution of "no feature-label dependence", and the achievable accuracy under the null is the label marginal $\max_c \pi_c$ plus finite-sample memorisation. Ojala and Garriga (2010), Permutation Tests for Studying Classifier Performance (JMLR), formalise both this null and a second variant (permuting within-feature) that tests feature-dependence structure. The training-set gap between real and permuted labels is exactly Zhang et al. (2017)'s random-label experiment used in the opposite direction: there, to show capacity; here, to certify signal.

Label noise: what mislabelling does at scale

  • Natarajan et al. (2013), Learning with Noisy Labels — risk-consistent surrogate losses under class-conditional noise rates.
  • Symmetric label noise at rate $\rho$ caps clean-test accuracy; deep nets fit the noise (memorisation) unless regularised or early-stopped — Arpit et al. (2017) document the patterns-before-noise dynamic that makes early stopping a partial noise defence.
  • Northcutt et al. (2021), Confident Learning: Estimating Uncertainty in Dataset Labels — principled detection of mislabelled examples from model confidence; the cleanlab implementation operationalises it. Their companion audit (Northcutt et al., 2021, Pervasive Label Errors in Test Sets) found roughly 3–6% label errors in canonical benchmarks including ImageNet — miscalibrating every leaderboard comparison that treats test labels as ground truth.

The out-of-bounds assert on CUDA

nll_loss kernels validate 0 <= target < C device-side; failure raises an asynchronous assert surfacing at a later synchronisation point with a stack trace pointing anywhere. CUDA_LAUNCH_BLOCKING=1 or a CPU rerun recovers the true site. The pre-emptive bincount/assert audit is cheaper than either — and unlike them, it also catches the legal-but-wrong encodings (unused class 0; a C+1-class head "fixing" 1-indexed labels) that raise nothing.

Alignment as a data-engineering invariant

The durable fix is structural: carry (feature_ref, label, sample_id) as one record through every transform, and checksum pairings at pipeline boundaries — the same provenance discipline as data versioning. Split by sample_id hash, not row order, so filtering or re-sorting cannot silently re-pair; hash-based splitting also keeps membership stable across dataset revisions, the property train/val reshuffles quietly violate.

What to learn next