GPUs: Memory, Scheduling and Cost
Tensor parallel serving
When a model is too big to fit on one GPU, you can split each individual layer across several GPUs, so they compute one answer together.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Tensor parallel serving splits a single layer of a model across several GPUs. They work together on one answer, instead of one GPU doing it alone.
The analogy you have already lived
A heavy, long log needs more than one person to lift. Four people each grab one end or one side, lift at the same moment, and carry it together. No single person could have moved it alone. Yet the log itself was never cut into pieces — it moved as one thing, carried by several people in coordination.
A very large model layer works the same way when it does not fit in one GPU's memory. Several GPUs each hold a slice of that one layer. They compute their part of the answer together, at the same moment.
Why it exists
Some models are too large for a single GPU's memory. A GPU with a fixed amount of memory cannot hold a model bigger than that memory. That stays true no matter how it is written or optimised.
Splitting the model layer by layer across GPUs is one answer. It makes each GPU wait for the one before it to finish though — a slow relay race. Tensor parallelism instead splits inside each layer. Every GPU works on the same layer at the same time, each holding and computing only its own slice.
How it works
ONE big layer, too large for one GPU:
[ ================ full weight matrix ================ ]
split into slices, one per GPU, computed AT THE SAME TIME:
GPU 0: [====] GPU 1: [====] GPU 2: [====]
| | |
computes its slice computes its slice computes its slice
of the answer of the answer of the answer
| | |
\___________________|____________________/
|
combine the slices into
ONE complete answerThe GPUs have to talk to each other to combine their slices. That conversation happens over a fast connection between them — for every single layer, every single request.
A real example you have seen
The very large language models behind the most capable chatbots are often too big to fit on a single GPU. However powerful that GPU is. Serving them at all requires exactly this kind of splitting, across several GPUs working in lockstep on every question.
The honest part
This is genuinely hard to see for yourself without a real multi-GPU machine. Most learners do not have one sitting on their desk. What follows is real, working code that shows the actual splitting mathematics correctly. It runs two "devices" as two halves of a normal Python program, not as two physical GPUs talking over real hardware. The idea it teaches is real. The setting it teaches it in is a simplification, made honestly, on purpose.
Remember this
- Tensor parallelism splits inside one layer, across GPUs that all work on it together.
- It exists because some models are too large to fit on a single GPU, no matter how well optimised.
- The GPUs must communicate on every layer, which is real, extra work multi-GPU serving has to pay for.
What to learn next
- FSDP: sharding a model across GPUs — a different way to split a model across GPUs, built for training rather than serving.
- Choosing a GPU for inference — deciding whether you need this at all, or whether one GPU is enough.
- Loading models bigger than your GPU — a simpler first option worth ruling out before reaching for tensor parallelism.
Developer — Code and libraries.
Setup
pip install torchSimulating the split, and proving it is correct
import torch
torch.manual_seed(0)
IN_FEATURES, OUT_FEATURES = 8, 6
x = torch.randn(1, IN_FEATURES)
W = torch.randn(OUT_FEATURES, IN_FEATURES) # one big weight matrix, as if on one GPU
b = torch.randn(OUT_FEATURES)
# ---- One device, the whole layer ----
full_output = x @ W.T + b
# ---- Two "devices": each holds HALF the output features (column-parallel) ----
split = OUT_FEATURES // 2
W_gpu0, b_gpu0 = W[:split], b[:split]
W_gpu1, b_gpu1 = W[split:], b[split:]
out_gpu0 = x @ W_gpu0.T + b_gpu0 # computed independently, "on GPU 0"
out_gpu1 = x @ W_gpu1.T + b_gpu1 # computed independently, "on GPU 1"
# The only communication needed: combine each device's slice of the answer.
combined_output = torch.cat([out_gpu0, out_gpu1], dim=-1)
print("full (one device) output: ", full_output.numpy().round(4))
print("combined (two devices) output:", combined_output.numpy().round(4))
print("outputs match exactly:", torch.allclose(full_output, combined_output))
print("GPU 0 held", W_gpu0.numel(), "of", W.numel(), "weight values")
print("GPU 1 held", W_gpu1.numel(), "of", W.numel(), "weight values")full (one device) output: [[ 0.7623 -1.6981 -1.9648 -5.1793 -3.8357 0.8765]] combined (two devices) output: [[ 0.7623 -1.6981 -1.9648 -5.1793 -3.8357 0.8765]] outputs match exactly: True GPU 0 held 24 of 48 weight values GPU 1 held 24 of 48 weight values
This is a real, exact numerical result, not a simplification of the maths. Splitting a linear layer's output features this way, then combining the results, genuinely produces the same answer as running the whole layer on one device. What this script does not show is the real cost of doing this across actual separate GPUs. The time spent physically moving data between them needs real multi-GPU hardware to measure.
Line-by-line walkthrough
Splitting W along its first dimension (W[:split], W[split:]). This is called column-parallelism — each device computes a different slice of the output, using the full input. Every device needs a full copy of the input x, but only a slice of the weights.
torch.cat([out_gpu0, out_gpu1], dim=-1). In a real multi-GPU setup, this step is not free. It is a genuine data transfer between GPUs, over a real physical connection like NVLink or PCIe. It happens on every single layer of a large model, for every single request.
torch.allclose(...). Confirms the split computation is mathematically identical to the unsplit one, floating-point rounding aside. This is the property that makes tensor parallelism trustworthy: it changes where computation happens, not what the computation is.
Common mistakes
Assuming more GPUs always means faster serving. Every extra GPU in a tensor-parallel group adds one more communication step per layer. Past a certain point, that communication cost outweighs the benefit of splitting further. See the researcher section for the exact trade-off.
Splitting a layer along the wrong dimension. The column-parallel split above needs no communication until the very end. Splitting along the input dimension instead, row-parallelism, needs a different combining step: summing partial results, not concatenating them. Mixing the two up produces confidently wrong numbers, not an error.
Forgetting the GPUs need a very fast connection to each other. This pattern assumes GPUs can exchange data with very low latency, many times per second. Running it across GPUs connected by a slow, ordinary network link turns the communication step into the dominant cost.
Try it yourself
Change OUT_FEATURES to 9 and split it unevenly (say, 5 and 4) between the two simulated devices. Confirm the combined output still exactly matches the full, unsplit computation.
What to learn next
- FSDP: sharding a model across GPUs — a different way to split a model across GPUs, built for training rather than serving.
- Choosing a GPU for inference — deciding whether you need this at all, or whether one GPU is enough.
- Loading models bigger than your GPU — a simpler first option worth ruling out before reaching for tensor parallelism.
Researcher — Mathematics and papers.
Column-parallel and row-parallel layers, paired
Megatron-LM (Shoeybi et al., 2019) introduced the pairing used in almost every subsequent tensor-parallel implementation: a column-parallel linear layer (splitting output features across $p$ devices, as in the demo) followed by a row-parallel linear layer (splitting input features, combined by summation rather than concatenation). For a transformer's MLP block specifically, this pairing means the intermediate activation between the two linear layers never needs to be communicated at all — only one all-reduce (a sum across all $p$ devices) is needed per MLP block, and one more per attention block, rather than a communication step after every individual matrix multiply.
Communication cost
For $p$ devices and hidden size $d$, each all-reduce over an activation of size $d$ (per token) using a ring all-reduce algorithm costs approximately
$$T_{\text{comm}} \approx 2(p-1)\frac{d}{p \cdot B}$$
where $B$ is the achievable inter-GPU bandwidth. Because this term does not shrink proportionally as $p$ grows — the $(p-1)/p$ factor approaches 1 — communication overhead grows with the number of devices in the tensor-parallel group, which is why tensor parallelism is generally kept within a single node (where GPUs share fast NVLink interconnects) rather than spread across nodes connected by slower networking, in contrast to pipeline and data parallelism strategies, which communicate far less often and tolerate slower links better.
Tensor parallelism at inference time specifically
Training and inference share the mechanism but differ in what dominates cost. During autoregressive decoding, each forward pass processes very few tokens — often exactly one, per sequence, per step. So the all-reduce communicates a small amount of data very frequently. Communication latency, not bandwidth, tends to dominate at small batch sizes. This is a large part of why serving frameworks — vLLM, NVIDIA TensorRT-LLM, Text Generation Inference — implement tensor parallelism as a first-class, tightly optimised serving configuration. It is commonly exposed as a tensor_parallel_size or tp_size setting, rather than reused training-oriented parallelism code.
Where tensor parallelism fits among the alternatives
- Tensor parallelism — splits inside a layer; low latency overhead per step, requires fast interconnect, typically limited to GPUs within one node.
- Pipeline parallelism — splits across layers, with different GPUs owning different layers; communicates far less, but introduces pipeline "bubbles" (idle time) unless carefully scheduled.
- Data parallelism — every GPU holds a full copy of the model and processes different data. No per-layer communication, but it needs enough memory for a full model copy per device — which tensor parallelism specifically avoids.
Large-scale systems typically combine all three (Narayanan et al., 2021), choosing the degree of each based on the interconnect topology actually available.
Papers
- Shoeybi et al., Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism, 2019 — arxiv.org/abs/1909.08053
- Narayanan et al., Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM, SC 2021 — arxiv.org/abs/2104.04473
What to learn next
- FSDP: sharding a model across GPUs — a different way to split a model across GPUs, built for training rather than serving.
- Choosing a GPU for inference — deciding whether you need this at all, or whether one GPU is enough.
- Loading models bigger than your GPU — a simpler first option worth ruling out before reaching for tensor parallelism.