Positions in images and video
A picture has rows and columns, not a single index. Flatten it naively and two neighbouring patches end up far apart, so vision models encode each axis separately.
- 15 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
An image has two directions, so it needs two position signals. One flat index puts neighbouring patches far apart.
Think of a chessboard photographed and then read out square by square, left to right, top to bottom. Square 8 is the top-right corner. Square 9 is the leftmost square of the next row.
By the numbers, 8 and 9 are neighbours. On the board they are at opposite ends. Meanwhile square 8 sits directly above square 16, and those numbers are eight apart.
Reading a grid as one long list scrambles what "near" means. That is the whole problem.
How a vision model sees a picture
A vision transformer chops an image into square patches, usually 16 by 16 pixels. Each patch becomes one token, the same way a word becomes one token in text.
A 224 by 224 image gives a 14 by 14 grid, so 196 patches. Those get flattened into a list and fed to a transformer, which has no idea the list came from a grid.
the grid the flat list
┌────┬────┬────┐
│ 0 │ 1 │ 2 │ 0 1 2 3 4 5 6 7 8
├────┼────┼────┤
│ 3 │ 4 │ 5 │ patch 2 and patch 3 are next to
├────┼────┼────┤ each other in the list, and in
│ 6 │ 7 │ 8 │ opposite corners of a row
└────┴────┴────┘The fix: give each axis its own signal
Split the position channels in half. Use one half to say which row the patch is in, and the other half to say which column.
Now two patches side by side share a row signal and differ slightly in column. Two patches stacked vertically share a column signal and differ slightly in row. Distance in the encoding matches distance on the picture.
The original vision transformer instead used a plain learned table, one row per patch position, and let the model discover the grid structure. That works, and it needs the model to spend capacity learning something you could have told it.
The resolution problem
Here is a wrinkle text does not have. Change the image size and the number of patches changes.
A model trained on 224-pixel images has a 14 by 14 table. Show it a 384-pixel image and you need 24 by 24. There is no row for patch 300.
The standard answer is to resize the position table the way you resize a picture, using two-dimensional interpolation. It is the same trick applied to the positions instead of the pixels. It works well for images, because neighbouring patches genuinely resemble each other.
Video adds a third direction
A video is a stack of images over time. Now a patch has three coordinates: which frame, which row, which column.
Modern vision-language models split the rotary dials three ways, one group per axis. Text uses all three identically, which makes it behave exactly like ordinary text. An image freezes the time axis and varies row and column. A video varies all three.
That design lets one model handle a paragraph, a photo and a clip with a single positional mechanism.
Where you have seen this
- Google Lens and reverse image search.
- A chat assistant reading a screenshot and telling you which button is where.
- Video summarisation that knows an event happened near the start.
- Medical scan tools reporting the location of a finding.
Remember this
- Flattening a grid into a list destroys what "nearby" means.
- Vision models split their position channels across the row and column axes.
- Changing image size changes the patch count, so position tables get interpolated.
- Video adds time as a third axis, handled the same way.
What to learn next
- Vision transformers — the architecture these positions feed.
- Vision language models — where text and image positions have to coexist.
- Rotary position embeddings (RoPE) — the one-dimensional version this extends.
Developer — Code and libraries.
Setup
pip install torch numpyWritten against PyTorch 2.5.1 and NumPy 1.26.4. Everything runs on CPU with no downloads.
Two-dimensional sine-cosine positions, and resizing a learned table
import numpy as np
import torch
import torch.nn.functional as F
np.set_printoptions(precision=3, suppress=True)
def sincos_1d(dim, pos, base=10000.0):
i = np.arange(dim // 2)
freq = 1.0 / (base ** (2 * i / dim))
ang = pos.reshape(-1, 1) * freq.reshape(1, -1)
return np.concatenate([np.sin(ang), np.cos(ang)], axis=1)
def sincos_2d(dim, grid_h, grid_w):
"""Half the channels encode the row, half encode the column."""
assert dim % 4 == 0, "needs to split into two halves, each with sin and cos"
rows = np.repeat(np.arange(grid_h), grid_w) # 0 0 0 1 1 1 2 2 2 ...
cols = np.tile(np.arange(grid_w), grid_h) # 0 1 2 0 1 2 0 1 2 ...
return np.concatenate([sincos_1d(dim // 2, rows),
sincos_1d(dim // 2, cols)], axis=1)
pe = sincos_2d(dim=16, grid_h=3, grid_w=3)
print("9 patches in a 3x3 grid, 16 channels each:", pe.shape)
print("\npatch (0,0) top-left :", pe[0])
print("patch (0,2) top-right :", pe[2])
print("patch (2,0) bottom-left :", pe[6])
def sim(a, b):
return float(a @ b / (np.linalg.norm(a) * np.linalg.norm(b)))
print("\ncosine similarity from the top-left patch:")
for name, j in [("its right neighbour (0,1)", 1), ("two to the right (0,2)", 2),
("the patch below (1,0)", 3), ("the far corner (2,2)", 8)]:
print(f" {name}: {sim(pe[0], pe[j]):.4f}")
# What a ViT does when the image size changes at inference.
torch.manual_seed(0)
D, OLD = 64, 14 # 14x14 = 196 patches
learned = torch.randn(1, OLD * OLD + 1, D) # +1 for the [CLS] token
cls, patches = learned[:, :1], learned[:, 1:]
NEW = 16 # a bigger image at test time
grid = patches.reshape(1, OLD, OLD, D).permute(0, 3, 1, 2)
resized = F.interpolate(grid, size=(NEW, NEW), mode="bicubic", align_corners=False)
resized = resized.permute(0, 2, 3, 1).reshape(1, NEW * NEW, D)
out = torch.cat([cls, resized], dim=1)
print(f"\nposition table for {OLD}x{OLD} patches: {tuple(learned.shape)}")
print(f"after 2-D interpolation to {NEW}x{NEW} : {tuple(out.shape)}")
print("the [CLS] row is copied, never interpolated:",
torch.equal(out[:, 0], learned[:, 0]))9 patches in a 3x3 grid, 16 channels each: (9, 16) patch (0,0) top-left : [0. 0. 0. 0. 1. 1. 1. 1. 0. 0. 0. 0. 1. 1. 1. 1.] patch (0,2) top-right : [ 0. 0. 0. 0. 1. 1. 1. 1. 0.909 0.199 0.02 0.002 -0.416 0.98 1. 1. ] patch (2,0) bottom-left : [ 0.909 0.199 0.02 0.002 -0.416 0.98 1. 1. 0. 0. 0. 0. 1. 1. 1. 1. ] cosine similarity from the top-left patch: its right neighbour (0,1): 0.9419 two to the right (0,2): 0.8205 the patch below (1,0): 0.9419 the far corner (2,2): 0.6409 position table for 14x14 patches: (1, 197, 64) after 2-D interpolation to 16x16 : (1, 257, 64) the [CLS] row is copied, never interpolated: True
Reading that output
The channel split is visible in the raw vectors. Patch (0,2) has zeros and ones in its first eight channels, identical to patch (0,0), because they share row 0. Its last eight channels differ, because the columns differ. Patch (2,0) is the mirror image: last eight identical, first eight changed.
The similarity to the right neighbour and to the patch below are both 0.9419. Exactly equal. One step right and one step down are the same distance, and the encoding says so. A flat index would rate them 1 apart and 3 apart.
The far corner scores 0.6409, the lowest of the four. Two steps in both directions is the furthest patch on the grid, and the encoding reflects that too.
The table went from 197 rows to 257. That is 196 patches plus one [CLS] token, resized to 256 patches plus one. The count is grid * grid + 1, and forgetting the +1 is the single most common bug in this code.
[CLS] is copied unchanged. It has no position on the grid, so interpolating it alongside the patches would be meaningless. Every correct implementation slices it off, interpolates the rest, and concatenates it back.
Two-dimensional RoPE
RoPE extends the same way. Split the head dimensions into two groups; rotate one group by the row index and the other by the column index.
import numpy as np
HEAD_DIM, BASE = 16, 100.0
half = HEAD_DIM // 2
inv = 1.0 / (BASE ** (np.arange(0, half, 2) / half)) # speeds for one axis
def rope_axis(x, pos, inv_freq):
ang = pos * inv_freq
c, s = np.cos(ang), np.sin(ang)
e, o = x[0::2], x[1::2]
out = np.empty_like(x)
out[0::2] = e * c - o * s
out[1::2] = e * s + o * c
return out
def rope_2d(x, row, col):
"""First half of the channels carries the row, second half the column."""
return np.concatenate([rope_axis(x[:half], row, inv),
rope_axis(x[half:], col, inv)])
rng = np.random.default_rng(0)
q, k = rng.normal(size=HEAD_DIM), rng.normal(size=HEAD_DIM)
print("score depends only on the row and column GAP, not on absolute place:")
for (r1, c1), (r2, c2) in [((0, 0), (1, 2)), ((3, 4), (4, 6)), ((10, 10), (11, 12))]:
print(f" q at ({r1},{c1}) k at ({r2},{c2}) gap (+1,+2):"
f" {rope_2d(q, r1, c1) @ rope_2d(k, r2, c2):.6f}")
print("\ndifferent gaps from the same starting patch (5,5):")
for dr, dc in [(0, 0), (0, 1), (1, 0), (1, 1), (3, 3)]:
print(f" gap ({dr},{dc}): {rope_2d(q, 5, 5) @ rope_2d(k, 5 + dr, 5 + dc):9.6f}")score depends only on the row and column GAP, not on absolute place: q at (0,0) k at (1,2) gap (+1,+2): 2.175659 q at (3,4) k at (4,6) gap (+1,+2): 2.175659 q at (10,10) k at (11,12) gap (+1,+2): 2.175659 different gaps from the same starting patch (5,5): gap (0,0): 2.472446 gap (0,1): 1.774216 gap (1,0): 2.397445 gap (1,1): 1.699215 gap (3,3): 2.810077
Three identical scores for the same two-dimensional offset at three different places on the grid. The relative property from ordinary RoPE holds per axis.
Two details worth noticing. The base is 100 rather than 10000, because a patch grid is tens of steps wide rather than thousands of tokens long, and the dials should match the range they have to cover. And the gap (3,3) scores higher than the gap (1,1), which is the same non-monotone behaviour ordinary RoPE has. RoPE makes the offset readable; it does not build in a preference for nearby patches.
Video, and the three-axis version
Qwen2-VL introduced M-RoPE, which splits the dials three ways: time, height and width. The rule for filling in the three position ids is what makes one model handle every modality:
| Input | temporal id | height id | width id |
|---|---|---|---|
| text token | t | t | t |
| image patch | constant per image | patch row | patch column |
| video patch | frame index | patch row | patch column |
When all three ids are equal, the three groups rotate by the same angle and M-RoPE reduces exactly to ordinary one-dimensional RoPE. Text behaves as it always did, at zero cost.
Common mistakes
Forgetting the [CLS] row when interpolating. The table has grid * grid + 1 rows. Reshaping all of them into a square fails, or worse, succeeds with an off-by-one that silently shifts every patch.
Using mode="nearest" for the interpolation. ViT position tables are smooth in the spatial axes. Bicubic is standard; nearest-neighbour duplicates rows and creates visible blocking in the positional signal.
Reusing the text RoPE base for a patch grid. A base of 10000 gives wavelengths up to 62832, wildly oversized for a 24-step axis. Almost all the dials become constant across the whole image.
Assuming rectangular images work automatically. A 14 by 20 grid needs different row and column handling. Code that assumes a square grid produces transposed positions on non-square inputs, and the model degrades in a way that looks like a data problem.
Try it yourself
Change sincos_2d to use a single flat index instead of splitting row and column, then reprint the four similarities. The right neighbour and the patch below will no longer match, and the far corner may score higher than a genuine neighbour. That broken table is what a naive flatten gives you.
What to learn next
- Vision transformers — the architecture these positions feed.
- Vision language models — where text and image positions have to coexist.
- Rotary position embeddings (RoPE) — the one-dimensional version this extends.
Researcher — Mathematics and papers.
The vision transformer baseline
Dosovitskiy et al. (2021) use a learned absolute embedding $E_{pos} \in \mathbb{R}^{(N+1) \times d}$ for $N = (H/P)(W/P)$ patches plus one class token, added to the patch projections. The paper's own ablation reports that 2-D-aware position embeddings gave no consistent benefit over the flat 1-D learned table at their scale, and attributes this to the model learning the grid structure from data. Later work at other scales and with other objectives does find 2-D structure useful, so this is not settled in one direction.
Fixed 2-D sine-cosine
The standard fixed construction, used in MAE and in DETR-family detectors, allocates $d/2$ channels to each axis:
$$ PE(r, c) = \big[\, \mathrm{sincos}{d/2}(r) \;\Vert\; \mathrm{sincos}{d/2}(c) \,\big] $$
where $\mathrm{sincos}_m(p)$ is the standard 1-D construction of width $m$. The concatenation makes the inner product separable:
$$ PE(r_1,c_1)^\top PE(r_2,c_2) = f(r_1 - r_2) + g(c_1 - c_2) $$
so grid distance decomposes additively across axes. That separability is the property the developer output verifies, and it is the reason one step right and one step down score identically.
Chu et al. (2023), CPVT, take a different route: generate positions with a depthwise convolution over the token grid, so the encoding is conditioned on content and resolution-agnostic by construction.
Resolution change
For a source grid $G \times G$ and target $G' \times G'$, treat $E_{pos}$ minus its class row as a $G \times G$ image with $d$ channels and resample bicubically. This is the standard adaptation for fine-tuning at higher resolution, used in ViT and DeiT.
Beyer et al. (2023), FlexiViT, train a single model to accept multiple patch sizes by resizing the patch-embedding kernel and position table during training, so the interpolation is in-distribution rather than a test-time hack. Dehghani et al. (2023), NaViT, take native-resolution training further with sequence packing, avoiding a fixed grid altogether.
RoPE in vision
Heo et al. (2024), Rotary Position Embedding for Vision Transformer, study axial 2-D RoPE, where head dimensions are partitioned by axis and rotated by $r$ and $c$ separately, and a mixed variant where learned frequencies span both axes. They report gains on classification, detection and segmentation, and particularly on resolution changes at test time, since RoPE requires no table to interpolate.
Axial 2-D RoPE gives
$$ \langle R_{(r_1,c_1)} q,\; R_{(r_2,c_2)} k \rangle = g\big(q, k,\, r_1 - r_2,\, c_1 - c_2\big) $$
with the relative property holding independently per axis, as the developer output shows numerically. Purely axial rotation cannot represent diagonal relative offsets as a single interaction, which is the motivation for the mixed-frequency variant.
M-RoPE for video and multimodal input
Wang et al. (2024), Qwen2-VL, partition the rotary dimensions into temporal, height and width groups. Position ids are assigned per modality as tabulated in the developer section. Two consequences follow.
Setting $t = h = w$ for text tokens makes M-RoPE numerically identical to 1-D RoPE on text, so no capability is traded away.
Because the maximum id used by an image or video block is bounded by its grid rather than by its token count, a long visual input consumes fewer positional steps than it does sequence steps. The Qwen2-VL report notes this helps extrapolation to longer videos than were seen in training.
Qwen2.5-VL extends the temporal axis to align with absolute timestamps rather than frame indices, so a model can reason about elapsed time rather than frame count, and handles variable frame rates.
Open questions
- Axial decomposition assumes row and column are independent. Learned mixed frequencies (Heo et al.) relax that and help. How much structure is genuinely two-dimensional rather than separable is unsettled.
- Native-resolution training (NaViT) and flexible patching (FlexiViT) reduce the interpolation problem but complicate batching. The tradeoff against fixed-grid training is workload dependent.
- For video, the right relative weighting between one frame of time and one patch of space is a hyperparameter with little theory behind it.
Papers
- Dosovitskiy et al., An Image is Worth 16x16 Words (ViT), ICLR 2021 — arxiv.org/abs/2010.11929
- He et al., Masked Autoencoders Are Scalable Vision Learners (MAE), CVPR 2022 — arxiv.org/abs/2111.06377
- Chu et al., Conditional Positional Encodings for Vision Transformers, ICLR 2023 — arxiv.org/abs/2102.10882
- Beyer et al., FlexiViT: One Model for All Patch Sizes, CVPR 2023 — arxiv.org/abs/2212.08013
- Dehghani et al., Patch n' Pack: NaViT, NeurIPS 2023 — arxiv.org/abs/2307.06304
- Heo et al., Rotary Position Embedding for Vision Transformer, ECCV 2024 — arxiv.org/abs/2403.13298
- Wang et al., Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution, 2024 — arxiv.org/abs/2409.12191
What to learn next
- Vision transformers — the architecture these positions feed.
- Vision language models — where text and image positions have to coexist.
- Rotary position embeddings (RoPE) — the one-dimensional version this extends.