PyTorch Tensors

reshape, view, permute and contiguity

view, reshape and permute rearrange how a tensor is read without moving its numbers — until contiguity forces a real copy, and an error tells you why.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A tensor's numbers sit in one long line in memory, and its shape is a set of instructions for reading that line.

Think of a bookshelf. The books stand in one long row, but you can read the row as "three shelves of four" or "four shelves of three" without moving a single book. Only your way of counting changes. That is what reshaping a tensor does.

Why it exists

Moving numbers around in memory is slow — millions of numbers, millions of moves. Changing the reading instructions is instant, no matter how big the tensor is.

So PyTorch keeps the numbers still whenever it can, and swaps the instructions instead. A reading of the same memory under new instructions is called a view.

How it works

memory:   [0, 1, 2, 3, 4, 5]        (never moves)

read as 2 rows of 3:      read as 3 rows of 2:
   0  1  2                   0  1
   3  4  5                   2  3
                             4  5

One line of memory, two different grids — same numbers.

There is a catch. Some rearrangements, like turning rows into columns, scramble the reading order so much that the old line cannot serve the new grid. Then PyTorch must copy the numbers into a fresh line after all. The word for "reading order still matches memory order" is contiguous.

A real example you have seen

A photo arrives from the internet as one long stream of bytes. Your phone displays it as height by width by colour without rearranging that stream — it reads the same bytes through a grid of instructions. Every image you have ever viewed was a "view".

Remember this

  • Numbers live in one long line; shape is instructions for reading it.
  • A view changes instructions, not data — that is why it is free.
  • Some rearrangements break the reading order and force a real copy.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

Outputs verified with torch 2.5.1, CPU.

Views share memory — watch it happen

views.py
import torch

x = torch.arange(6)          # [0, 1, 2, 3, 4, 5] in one block of memory
v = x.view(2, 3)             # same memory, read as 2 rows of 3
v[0, 0] = 99                 # writing through the view...
print(x)                     # ...changes the original

t = v.t()                    # transpose: only the reading order changed
print(t.is_contiguous())     # the reading order no longer matches memory

print(v.stride(), t.stride())

print(t.reshape(6))              # reshape copies when it has to
print(t.contiguous().view(6))    # or: repack memory first, then view
Output
tensor([99,  1,  2,  3,  4,  5])
False
(3, 1) (1, 3)
tensor([99,  3,  1,  4,  2,  5])
tensor([99,  3,  1,  4,  2,  5])

The walkthrough

v[0, 0] = 99 changed x. One block of memory, two readers. This is the single most important fact on this page: view is not a copy, and editing a view edits the original.

stride() is the reading instructions made visible. Stride (3, 1) means: move 3 places in memory to go down a row, 1 place to go right a column. The transpose swapped it to (1, 3) — same memory, opposite walking pattern.

t.is_contiguous() is False. Walk t row by row and you visit memory as 99, 3, 1, 4, 2, 5 — out of order. view refuses such tensors, because a view can only re-slice memory that is already in reading order:

python
torch.arange(6).view(2, 3).t().view(6)
Output
RuntimeError: view size is not compatible with input tensor's size and stride (at least one dimension spans across two contiguous subspaces). Use .reshape(...) instead.

reshape is the polite version. It returns a free view when memory order allows, and silently copies when it does not. contiguous() is the explicit version: repack memory now, then everything downstream is view-cheap.

permute, the many-dimensional transpose

t() only swaps two dimensions. permute reorders any number of them — the classic case is image formats, channels-first versus channels-last:

python
img = torch.rand(3, 32, 48)          # channels, height, width
hwc = img.permute(1, 2, 0)           # height, width, channels
print(hwc.shape, hwc.is_contiguous())
Output
torch.Size([32, 48, 3]) False

Like transpose, permute never moves data — so its result is usually non-contiguous, and a following view will need contiguous() first.

Common mistakes

Using view to "fix" a shape error. If the shapes disagree, view will happily give you a tensor of the right shape and the wrong meaning — rows sliced mid-sample, images woven together. A network fed this trains badly instead of crashing. Reshape only when you can say where each number should land.

Confusing permute with reshape. x.permute(1, 0) and x.reshape(cols, rows) produce the same shape and different tensors. Permute reorders the reading; reshape re-slices the line. Print a small example of both once and the difference sticks.

Assuming reshape never copies. It copies exactly when the input is non-contiguous. Free in the demo, 100 MB in production. If you need a guarantee, call view — the error is the guarantee.

Editing what you thought was a copy. Slices and views alias the original. When you need independence, say .clone().

Try it yourself

Create torch.arange(24).view(2, 3, 4), then permute(2, 0, 1), and predict the stride before printing it. Then predict which of .view(24) and .reshape(24) succeeds, and check.

What to learn next

Researcher — Mathematics and papers.

The strided memory model

A tensor is (data pointer, shape, strides, storage offset). Element at index (i_0, …, i_{k−1}) lives at:

addr = offset + Σ_d i_d × stride_d

Where stride_d is the memory step (in elements, not bytes) for dimension d, and k is the rank. A tensor is contiguous (row-major) when stride_{k−1} = 1 and stride_d = stride_{d+1} × shape_{d+1} — the strides a fresh allocation would get.

Every "free" operation — view, transpose, permute, expand, basic slicing, narrow — is arithmetic on this quadruple with the data pointer unchanged. view(new_shape) succeeds iff the new shape's implied traversal can be expressed as strides over the existing layout; the compatibility condition (each new dimension must map onto a contiguous run of old dimensions) is checked in computeStride, and failure is the error shown above.

Costs

  • View/permute/slice: O(1) time and memory, always.
  • contiguous() / copying reshape: O(n) — a full traversal in logical order, gathering scattered elements. On GPU this is bandwidth-bound; on non-contiguous layouts it also defeats coalesced memory access, so the copy can cost more per element than a contiguous one.
  • Downstream cost of ignoring layout: many kernels have fast paths for contiguous inputs and fall back to strided element-wise indexing otherwise. One contiguous() before a hot loop can beat thousands of strided reads inside it.

channels_last and layout as an optimisation axis

The same logical (N, C, H, W) tensor can be stored NCHW-contiguous or NHWC-contiguous (memory_format=torch.channels_last), and convolution kernels on tensor-core hardware strongly prefer the latter. PyTorch tracks memory format through operations, so layout becomes a per-model switch rather than a code rewrite. The mechanism is the stride machinery on this page — nothing more exotic.

Lineage

The strided ndarray model comes from NumPy (Harris et al., 2020, Nature 585), itself descending from Numeric (1995). PyTorch inherited it via Torch7's Tensor/Storage split (Collobert et al., 2011, Torch7: A Matlab-like Environment for Machine Learning). The design's endurance is the point: one abstraction supports views, broadcasting (stride 0 — see broadcasting), and layout optimisation without new machinery.

What to learn next