Debugging PyTorch

Expected scalar type Float but found Double

Dtype errors come from three habitual sources — NumPy's float64, integer labels in float losses, and float labels in integer losses — and each has a one-line cast as its fix.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have seen this
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

This error means two tensors met an operation while storing their numbers in different formats — and PyTorch wants you to pick one format on purpose.

It is a currency problem. Someone hands the shopkeeper a fifty-rupee note and a five-dollar bill for one payment. Both are real money. The shopkeeper still refuses the pile, because mixing currencies in one transaction invites silent mistakes about value.

A tensor's number format is called its dtype (data type) — how much memory each number takes, and whether it can hold decimals. Mixing formats is the mixed pile of notes.

Why it exists

Different formats make different trades. A precise format spends more memory per number and computes slower. A lighter format is fast and small but rounds more aggressively. Whole-number formats cannot hold decimals at all.

PyTorch could convert secretly. For genuinely mixed pairs it refuses instead, because each secret conversion would either waste speed or quietly lose precision — and you would never know which. The refusal makes you choose.

How it works

model's numbers:  standard format          your data: extra-precise format
                    └──────── ✗ refused ───────┘
                    fix: convert once, where data enters

Nearly always, one side arrived from outside — a file, a spreadsheet, another library — carrying a foreign format. Convert at the border, once, and everything downstream agrees.

Where you have seen this

Same trap as international travel: you exchange money once at the airport, not at every shop. Data entering a model deserves the same treatment — one exchange at the border, not scattered conversions everywhere.

Remember this

  • dtype = the storage format of a tensor's numbers.
  • Mixing formats in one operation triggers the error; PyTorch will not guess.
  • Convert once, at the border where outside data comes in.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch numpy

Captured with torch 2.5.1 on CPU. This error wears several outfits — the exact words depend on which operation caught the mismatch, so this lesson shows the three you will actually meet.

Outfit one: NumPy's float64 meets the model

NumPy's default decimal format is float64 ("Double"); PyTorch models default to float32 ("Float"). Any data arriving through NumPy — CSVs via pandas, .npy files, sklearn arrays — carries the foreign dtype in:

dtype_crash_numpy.py
import numpy as np
import torch
import torch.nn as nn

torch.manual_seed(0)                         # so the printed weights reproduce

data = np.array([[0.5, 1.2], [0.3, 0.8]])    # NumPy defaults to float64
x = torch.from_numpy(data)                   # the tensor inherits float64
print("tensor dtype:", x.dtype)

model = nn.Linear(2, 1)                      # weights are float32
out = model(x)
Output
tensor dtype: torch.float64
RuntimeError: mat1 and mat2 must have the same dtype, but got Double and Float

Outfit two: integer labels in a float loss

dtype_crash_mse.py
import torch
import torch.nn as nn

pred = torch.randn(4, requires_grad=True)
target = torch.tensor([1, 0, 1, 0])          # integers: dtype int64
loss = nn.MSELoss()(pred, target)
loss.backward()
Output
RuntimeError: Found dtype Long but expected Float

MSELoss is regression machinery; it wants decimal targets. Integer literals produce int64 ("Long") tensors.

Outfit three: float labels in an integer loss

dtype_crash_ce.py
import torch
import torch.nn as nn

logits = torch.randn(4, 3)
target = torch.tensor([2.0, 0.0, 1.0, 2.0])  # class labels stored as floats
loss = nn.CrossEntropyLoss()(logits, target)
Output
RuntimeError: expected scalar type Long but found Float

The mirror image: CrossEntropyLoss wants class indices, which must be Long. (Float targets of shape (N, C) are legal as full probability distributions — a different feature; shape (N,) floats are the error.)

The fixes, all together

dtype_fixed.py
import numpy as np
import torch
import torch.nn as nn

torch.manual_seed(0)                         # so the printed weights reproduce

data = np.array([[0.5, 1.2], [0.3, 0.8]])
x = torch.from_numpy(data).float()           # cast once, at the border
print("input dtype:", x.dtype)

model = nn.Linear(2, 1)
print("output:", model(x).detach().numpy().round(3))

mse_target = torch.tensor([1, 0], dtype=torch.float32)   # regression targets: float
ce_target = torch.tensor([1, 0], dtype=torch.long)       # class labels: long
print("mse target:", mse_target.dtype, "| ce target:", ce_target.dtype)
Output
input dtype: torch.float32
output: [[-0.129]
 [-0.28 ]]
mse target: torch.float32 | ce target: torch.int64

The decision table worth memorising:

TensorCorrect dtype
Model inputs (features)float32
Regression targets, BCE targetsfloat32
Class-index targets for CrossEntropyLoss / NLLLosslong (int64)
Indices for embeddings and gatherlong

Common mistakes

Fixing outfit one by converting the model. model.double() runs — and doubles memory while significantly slowing GPU compute, since consumer cards execute float64 at a tiny fraction of float32 speed. Convert the data down, not the model up.

Casting the same tensor in five places. Scattered .float() calls mean the day one is forgotten, the error returns. One conversion in the Dataset's __getitem__ (or wherever data enters) is the maintainable version — the natural home is a custom Dataset class.

.float() on class labels to silence outfit three. The error flips into outfit two's territory, or worse, trains against nonsense. Read which dtype the operation wanted; the fix direction differs each time.

Forgetting from_numpy shares memory. torch.from_numpy(data) wraps the same buffer — while .float() makes the needed copy anyway, dtype-preserving borderline cases can alias surprisingly. Tensor creation and dtypes covers the sharing rules.

Try it yourself

Predict the dtype of each: torch.tensor([1, 2]), torch.tensor([1.0, 2]), torch.arange(5), torch.ones(3), and torch.from_numpy(np.arange(3)). Print all five and score yourself — the mixed-list case surprises most people.

What to learn next

Researcher — Mathematics and papers.

Type promotion: where mixing IS allowed

Elementwise ops do promote across dtypes (float32 + int64 → float32; float32 + float64 → float64) following torch.result_type semantics, largely NumPy-compatible with the notable exception that Python scalars promote weakly. The hard errors of this lesson come from ops that opt out of promotion — matmul/linear kernels and loss functions with semantic dtype contracts — where a silent upcast would respectively cost a kernel specialisation or hide a labelling bug. Knowing which regime an op belongs to explains the otherwise confusing "adding worked, multiplying crashed" experience.

Why float32 is the default

  • Memory and bandwidth: half of float64, which on bandwidth-bound workloads is a direct 2x.
  • Hardware: consumer NVIDIA GPUs execute fp64 at 1/32 to 1/64 of fp32 throughput; only datacentre parts (A100/H100) carry serious fp64 units. Apple MPS lacks float64 entirely.
  • Statistical fit: gradient noise dwarfs fp32 rounding for neural training; fp64's extra 29 bits of mantissa buy nothing measurable in accuracy.

The direction of travel is downward: fp16/bf16 mixed precision as routine (torch.autocast handling per-op dtype policy automatically — its own curated version of this lesson), and fp8 formats on H100-class hardware (Micikevicius et al., 2022, FP8 Formats for Deep Learning). bf16's trade — fp32's exponent range with 8 mantissa bits — eliminates the loss-scaling apparatus fp16 needs, which reconnects to NaN debugging.

The Long requirement for indices

Index tensors (CE targets, Embedding inputs, gather indices) must be int64 in eager PyTorch — a fixed kernel contract predating int32 index support discussions; the practical consequence is that int32 labels from Arrow/Parquet pipelines need an explicit .long(). float64 creeping in via NumPy, and int32 creeping in via Arrow, are the two border-crossing failure modes worth codifying in loader code review.

Precision as a correctness tool

Upcasting is still a legitimate diagnostic: running a numerically suspicious module in float64 (module.double()) isolates whether a discrepancy is rounding or logic — the technique behind gradcheck, which runs in double precision for exactly this reason.

What to learn next