einsum for tensor operations
einsum is one function that expresses dot products, matrix multiplies and attention in a short label string — you name the dimensions, it does the rest.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
einsum is one function that can do most tensor maths, controlled by a short string of letters naming each direction.
Think of addressing a parcel. "Flat 3, Building B, Sector 7" — each part of the address names one level of where things live. einsum asks you to address your tensors the same way: give every direction a letter, like "b for batch, i for rows". Then you write a tiny recipe with those letters, and einsum follows it.
Why it exists
Tensor code fills up with matmuls, transposes, unsqueezes and sums, and after three of them nobody can say what shape anything is. The intent drowns in plumbing.
einsum flips it. You state what each direction is and what should happen to it. The recipe replaces the plumbing — and doubles as documentation, because the letters spell out your shapes.
How it works
A recipe looks like "ik,kj->ij". Left of ->: one letter-group per input. Right: the letters you want to keep.
Two rules run the whole machine:
same letter in two inputs -> those directions are matched up
letter missing after -> -> that direction is summed away"ik,kj->ij" k appears in both, and is gone after the arrow:
match along k, multiply, add it away.
That is matrix multiplication - said in seven characters.A real example you have seen
Every chatbot reply involves attention — each word deciding which other words matter. Inside the code of GPT-style models, that step is often literally one einsum recipe matching "every query against every key". One readable line, doing billions of multiplications.
Remember this
- Letters name directions; the recipe says which match and which stay.
- A letter missing after
->gets summed away. - One function covers dot products, matmuls, transposes and more.
What to learn next
- Moving tensors between CPU and GPU — where these operations pay off at speed.
- Shapes and broadcasting — the implicit rules einsum makes explicit.
- Transformers — the architecture whose core is two einsums.
Developer — Code and libraries.
Setup
pip install torchOutputs verified with torch 2.5.1, CPU.
Five recipes that cover most real uses
import torch
v = torch.tensor([1., 2., 3.])
w = torch.tensor([10., 20., 30.])
print(torch.einsum("i,i->", v, w)) # dot product: multiply, then sum away i
print(torch.einsum("i,j->ij", v, w)) # outer product: nothing summed away
A = torch.arange(6.).reshape(2, 3)
B = torch.arange(12.).reshape(3, 4)
C = torch.einsum("ik,kj->ij", A, B) # matrix multiply: k is summed away
print(torch.allclose(C, A @ B))
batch_a = torch.ones(5, 2, 3)
batch_b = torch.ones(5, 3, 4)
print(torch.einsum("bik,bkj->bij", batch_a, batch_b).shape)
sq = torch.arange(9.).reshape(3, 3)
print(torch.einsum("ii->", sq)) # trace: sum of the diagonaltensor(140.)
tensor([[10., 20., 30.],
[20., 40., 60.],
[30., 60., 90.]])
True
torch.Size([5, 2, 4])
tensor(12.)The walkthrough
Dot versus outer is the whole grammar in two lines. Same letter (i,i->): positions pair up, then i vanishes — multiply and total. Different letters (i,j->ij): nothing pairs, nothing vanishes — every combination survives. Once you can predict those two, you can read any recipe.
The batch recipe bik,bkj->bij is where einsum starts beating the alternatives. b appears everywhere: kept, matched, looped over — a matmul per batch item. No bmm, no unsqueeze dance. Read it aloud: "for each b, multiply i-by-k with k-by-j."
ii-> uses one letter twice in a single input: that walks the diagonal, then sums it — the trace. (ii->i would keep the diagonal itself.)
Attention in one line, the reason many people learn einsum at all:
q = torch.ones(2, 5, 8) # batch, queries, features
k = torch.ones(2, 7, 8) # batch, keys, features
print(torch.einsum("bqd,bkd->bqk", q, k).shape) # every query vs every keytorch.Size([2, 5, 7])
d (features) is matched and summed away; q and k both survive, giving the score grid. The recipe even documents the tensors' meanings.
Common mistakes
Expecting element-wise, writing matmul — or the reverse. "ij,ij->ij" is element-wise multiply (all letters kept). "ij,jk->ik" is matmul. One letter of difference, entirely different maths, both run without error. Check which letters vanish.
A letter meaning two different sizes. If i is size 2 in one input and 3 in another, einsum raises a size-mismatch error — annoying but honest. The dishonest version is when both sizes coincide accidentally and a wrong recipe runs anyway. Name letters after meanings (b, q, k, d), not habit (i, j, k), and this mostly disappears.
Forgetting the output side. "ij->" sums everything; "ij->i" sums away only j. Omitting the arrow entirely triggers implicit mode with alphabetical output ordering — a NumPy inheritance that surprises. Always write the arrow.
Dismissing einsum as slow. For two operands, torch lowers einsum to the same matmul kernels you would call yourself. For 3+ operands, contraction order matters enormously — see the researcher block.
Try it yourself
Write recipes for: transposing a matrix; summing a (5, 3) tensor per-column; and a weighted average over keys — "bqk,bkd->bqd" given scores and values. Predict every output shape before running, then check with .shape.
What to learn next
- Moving tensors between CPU and GPU — where these operations pay off at speed.
- Shapes and broadcasting — the implicit rules einsum makes explicit.
- Transformers — the architecture whose core is two einsums.
Researcher — Mathematics and papers.
The notation, formally
einsum evaluates a sum-of-products over all index assignments. For "ik,kj->ij":
out[i,j] = Σ_k A[i,k] · B[k,j]
Generally: free indices (those after ->) enumerate the output; bound indices (missing from the output) are summed over their full range. This is Einstein summation with explicit output specification — repeated-index-implies-summation, from tensor calculus (Einstein, 1916), restricted to the string alphabet and without covariant/contravariant distinction.
Any single einsum is a tensor contraction possibly composed with diagonal extraction (repeated letters within one operand) and broadcasting (... ellipsis for unnamed leading dimensions).
Complexity and contraction order
A single contraction over inputs with free-index sizes F and bound sizes S costs O(prod(F) · prod(S)) multiply-adds — for matmul, the familiar O(n·m·k).
With three or more operands, the pairwise contraction order changes cost by orders of magnitude — the matrix-chain problem generalised, and finding the optimal order for arbitrary networks is NP-hard (Chi-Chung, Sadayappan, Wenger, 1997). The opt_einsum library (Smith and Gray, 2018, opt_einsum — A Python package for optimizing contraction order, JOSS) provides greedy and dynamic-programming path optimisers; PyTorch integrates it via torch.backends.opt_einsum, enabled when the package is installed. For hot many-operand contractions, precomputing a path with opt_einsum.contract_expression and reusing it beats re-planning per call.
For two operands, torch reduces einsum to permutes + reshapes + bmm, so einsum-vs-matmul performance differences are noise; benchmark before rewriting either way.
einsum as a lens on attention
Scaled dot-product attention is three contractions:
scores = QK^T/√d : "bhqd,bhkd->bhqk" — cost O(b·h·q·k·d)
weights = softmax over k (not a contraction)
out = weights·V : "bhqk,bhkd->bhqd" — cost O(b·h·q·k·d)
Where b is batch, h heads, q/k sequence positions, d head dimension. The q·k factor is the quadratic attention cost; einsum notation makes it visible by inspection. Fused implementations (FlashAttention, Dao et al., 2022) compute the same contraction tiled in SRAM without materialising the (q, k) matrix — the algebra stays einsum, the schedule changes.
Successors
einops (Rogozhnikov, 2022, ICLR, Einops: Clear and Reliable Tensor Manipulations) extends the labelling idea to reshapes/rearrangements (rearrange, reduce) with named, sized axes, and is now common in research code alongside einsum. Named tensor proposals aim to move the labels into the tensor type itself; as of torch 2.x they remain prototype.
What to learn next
- Moving tensors between CPU and GPU — where these operations pay off at speed.
- Shapes and broadcasting — the implicit rules einsum makes explicit.
- Transformers — the architecture whose core is two einsums.