PyTorch Tensors

Creating tensors and choosing a dtype

Every tensor has a dtype — the size of box each number lives in — and most silent numerical bugs come from letting PyTorch choose it for you.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A tensor is a grid of numbers, and its dtype is the size of box each number is stored in.

Think of storing dal in kitchen containers. A small dabba is fine for a spoonful, but pour a full packet in and it overflows onto the counter. Numbers work the same way in a computer. Each one lives in a container of fixed size, and a value too big for its container spills.

The container size has a name: the dtype, short for data type. Every tensor has exactly one dtype, shared by all its numbers.

Why it exists

Computers store numbers in fixed amounts of memory. Bigger containers hold bigger, more precise numbers — but a million numbers in big containers is a lot of memory.

So PyTorch lets you choose. Precise and heavy, or rough and light. Deep learning almost always picks a middle option called float32: decimal numbers stored in 32 bits, roughly 7 reliable digits each.

How it works

your numbers  →  [ tensor ]  →  dtype decides:
                                 how much memory each number takes
                                 how big a value fits
                                 how many digits survive

When you create a tensor, PyTorch looks at your numbers and picks a dtype. Decimals become float32. Whole numbers become int64 — a large container for whole numbers only, no decimal point.

A real example you have seen

Phone assistants like Google Translate run models in smaller containers than they were trained in. Shrinking the dtype halves the memory again and again, until the model fits in your pocket. Done carelessly, the answers get worse — which is why dtype is a real decision, not a detail.

Remember this

  • Every tensor has one dtype: the container size for its numbers.
  • PyTorch guesses from your input: decimals → float32, whole numbers → int64.
  • Too big for the container means spilled — and spills can be silent.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

Everything here runs on CPU in under a second. Outputs verified with torch 2.5.1.

The creation toolkit

dtypes.py
import torch

decimals = torch.tensor([1.5, 2.5])   # decimal numbers -> float32
whole = torch.tensor([1, 2])          # whole numbers  -> int64
print(decimals.dtype, whole.dtype)

print(torch.zeros(2, 3))
print(torch.arange(0, 10, 2))
print(torch.linspace(0, 1, 5))

# mixing dtypes: PyTorch promotes to the type that loses less
print((decimals + whole).dtype)

# a dtype too small for the value wraps around with no error
print(torch.tensor([300]).to(torch.uint8))

# float16 keeps roughly 3 decimal digits
wide = torch.tensor([0.1], dtype=torch.float64)
narrow = wide.to(torch.float16)
print(wide.item(), "->", narrow.item())
Output
torch.float32 torch.int64
tensor([[0., 0., 0.],
        [0., 0., 0.]])
tensor([0, 2, 4, 6, 8])
tensor([0.0000, 0.2500, 0.5000, 0.7500, 1.0000])
torch.float32
tensor([44], dtype=torch.uint8)
0.1 -> 0.0999755859375

Line by line, the parts that bite

torch.tensor(...) always copies its input. Hand it a NumPy array and you get an independent copy. If you want sharing instead, that is torch.from_numpy, and it is a different decision.

.to(torch.uint8) wrapped 300 into 44. uint8 holds 0 to 255, so 300 wraps around: 300 − 256 = 44. No warning, no error. Interestingly, building it directly — torch.tensor([300], dtype=torch.uint8) — is caught, and raises value cannot be converted to type uint8 without overflow. Construction checks; casting does not.

float16 turned 0.1 into 0.0999755859375. Half precision has about 3 reliable decimal digits. Fine for a model's activations, dangerous for a running sum of millions of them.

.item() pulls a single Python number out of a one-element tensor. On anything bigger it raises a Tensor with 2 elements cannot be converted to Scalar.

Common mistakes

Integer division surprises. torch.tensor([1, 2]) / 2 gives floats — but older habits from other languages expect truncation. When you want floor division, say so: //.

Building a model input as int64. Layers expect float input. Linear fed an integer tensor raises a dtype error. Fix at creation: torch.tensor(data, dtype=torch.float32).

Accumulating in float16. Summing thousands of half-precision numbers drifts badly. Accumulate in float32, cast down at the end — this is what mixed precision training does under the hood.

Assuming torch.arange(0, 1, 0.1) behaves. Float steps accumulate rounding error and the last element becomes unpredictable. For decimal ranges, use linspace, which takes a count instead of a step.

Try it yourself

Predict the dtype of torch.tensor([True, False]) + torch.tensor([1, 2]), then run it. Then find the largest whole number float32 can store without skipping — start testing around 16 million, with torch.tensor(16_777_216.) + 1.

What to learn next

Researcher — Mathematics and papers.

Floating-point layout

An IEEE 754 float is sign, exponent and mantissa: value = (−1)^s × 1.m × 2^(e − bias).

dtypebitsexponentmantissa~decimal digitsmax value
float6464115215.9~1.8e308
float32328237.2~3.4e38
float16165103.365504
bfloat1616872.4~3.4e38

Where s is the sign bit, e the exponent field, m the mantissa bits and bias = 2^(bits_e−1) − 1.

The float16 versus bfloat16 trade is the practically important one. bfloat16 keeps float32's exponent — the same range — and pays with mantissa: less precision, but overflow effectively disappears. That is why bf16 training rarely needs loss scaling while fp16 training does (Micikevicius et al., 2017).

Integers above 2^mantissa_bits+1 are not all representable: float32 skips odd integers past 16,777,216. Any pipeline that stores IDs in floats is quietly corrupting them past that point.

Promotion semantics

PyTorch follows NumPy-style type promotion with one deliberate exception: Python scalars are "weak" — float32_tensor * 2.0 stays float32 rather than promoting to float64. Zero-dimensional tensors participate as weakly as scalars. The full lattice is documented under "type promotion" in the docs; the practical rule is that the wider category (bool < int < float < complex) wins, then the wider width within it.

Cost

Memory is exactly numel() × element_size() bytes plus constant overhead — a 7B-parameter model is 28 GB in float32, 14 GB in bf16, and ~3.5 GB at 4-bit quantisation. On modern accelerators, low-precision matmul is not only smaller but faster: tensor cores execute fp16/bf16 at multiples of fp32 throughput, which is why dtype is a speed decision, not only a memory one.

References

  • IEEE 754-2019, Standard for Floating-Point Arithmetic.
  • Micikevicius et al. (2017), Mixed Precision Training — fp16 with fp32 master weights and loss scaling.
  • Kalamkar et al. (2019), A Study of BFLOAT16 for Deep Learning Training — the case for trading mantissa for exponent.
  • Dettmers et al. (2022), LLM.int8() — where 8-bit inference breaks and the outlier-channel fix.

What to learn next