Expert capacity and dropped tokens
Each expert gets a fixed number of seats per batch, so an oversubscribed expert silently skips the extra tokens, and the size of that buffer is a direct trade between waste and dropping.
- 12 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Every expert has a fixed number of seats per batch. Tokens that arrive after the seats fill up are skipped.
The bus with fixed seats
A bus has fifty seats. Sixty people are waiting at the stop. Ten of them are not travelling on that bus, however much they want to.
You could send a hundred-seat bus instead. Now nobody is left behind, and half the seats travel empty, burning the same fuel.
That is the entire problem in this lesson. Buses that are too small leave people behind. Buses that are too big waste fuel. Neither is free.
Why the seats are fixed at all
A graphics card is fast because it does the same operation on a fixed-size block of numbers. It hates blocks that change size from one batch to the next.
So the software reserves a fixed buffer per expert before it knows how many tokens will arrive. That buffer size is the number of seats.
The size is set by a dial called the capacity factor. A factor of one means seats for exactly the average share. A factor of two means seats for double the average.
What happens to a skipped token
Nothing dramatic, and that is what makes this hard to notice.
The layer is skipped for that token. Its value passes through untouched, as if that layer were not there. The model keeps running and produces text that looks fine.
token gets a seat: value -> expert -> value + expert's contribution
token is skipped: value ------------> value (unchanged)No crash, no warning. Quality falls a little, quietly, in a way you find only if you measure.
The connection to the previous lesson
If traffic were perfectly even, every expert would need exactly its average share of seats and nothing would be dropped.
Traffic is never perfectly even. So the seat count has to cover the busiest expert, not the average one. Every other expert then travels half empty.
Uneven traffic and wasted seats are the same problem seen from two sides. That is why balancing matters so much.
The modern answer
Newer training systems removed the fixed seats entirely. They arrange the arithmetic so it handles ragged, uneven groups without padding.
No seats, no drops, no empty seats. Where that machinery is available, this whole dial disappears.
It has not disappeared everywhere. When experts are spread across many machines, fixed-size buffers are still the simplest way to plan network traffic. So seats are still common at serving time.
Remember this
- Each expert has a fixed buffer per batch, sized by the capacity factor.
- Tokens beyond the buffer are skipped, and pass through the layer unchanged and unannounced.
- Bigger buffers waste compute; smaller buffers drop tokens. Newer systems avoid the choice.
What to learn next
- Shared and fine-grained experts — smaller experts, and what that does to load statistics.
- Serving a MoE across GPUs — where fixed buffers still earn their keep.
- Load balancing and expert collapse — the upstream cause of every drop.
Developer — Code and libraries.
Setup
pip install torchWritten against PyTorch 2.5.1, Python 3.10. CPU, a few seconds.
Measuring drops against buffer size
import math, torch
torch.manual_seed(0)
T, E, K = 4096, 16, 2 # tokens in the batch, experts, top-k
# A realistic router: mildly unbalanced, not collapsed.
pref = torch.softmax(torch.randn(E) * 0.6, dim=0)
choice = torch.multinomial(pref.expand(T, E), K, replacement=False)
counts = torch.bincount(choice.flatten(), minlength=E)
print("tokens assigned per expert (top-2 routing, 4096 tokens):")
print(" ", counts.tolist())
print(f" mean {counts.float().mean():.0f} max {counts.max()} "
f"max/mean {counts.max()/counts.float().mean():.2f}")
def simulate(cf):
cap = math.ceil(T * K / E * cf) # buffer slots per expert
seen = torch.zeros(E, dtype=torch.long)
dropped_slots = 0
lost_entirely = 0
for t in range(T): # first come, first served, in token order
got_any = False
for e in choice[t].tolist():
if seen[e] < cap:
seen[e] += 1
got_any = True
else:
dropped_slots += 1
if not got_any:
lost_entirely += 1
return cap, dropped_slots, lost_entirely
print(f"\n{'capacity factor':>16}{'slots/expert':>14}{'buffer slots':>14}"
f"{'dropped':>10}{'drop rate':>11}{'tokens with no expert':>23}")
for cf in (0.5, 1.0, 1.25, 1.5, 2.0):
cap, dropped, lost = simulate(cf)
print(f"{cf:>16.2f}{cap:>14}{cap*E:>14,}{dropped:>10,}"
f"{dropped/(T*K):>11.2%}{lost:>23,}")
print("\nwhat the buffer costs: padded compute is capacity, not actual load")
for cf in (1.0, 1.25, 2.0):
cap = math.ceil(T * K / E * cf)
print(f" cf={cf:<5} buffer holds {cap*E:>7,} slots for {T*K:,} assignments "
f"-> {(T*K)/(cap*E):.0%} of the buffer is real work")
# A dropped token is not zeroed. The residual connection carries it through.
x = torch.tensor([1.0, 2.0, 3.0])
expert_out = torch.tensor([0.5, -0.5, 0.1])
print("\nwhat a dropped token actually experiences")
print(" processed:", (x + expert_out).tolist())
print(" dropped :", (x + torch.zeros(3)).tolist(), " <- the layer becomes the identity")tokens assigned per expert (top-2 routing, 4096 tokens):
[253, 274, 446, 373, 815, 748, 388, 127, 629, 236, 526, 576, 521, 971, 919, 390]
mean 512 max 971 max/mean 1.90
capacity factor slots/expert buffer slots dropped drop rate tokens with no expert
0.50 256 4,096 4,248 51.86% 1,633
1.00 512 8,192 1,609 19.64% 318
1.25 640 10,240 893 10.90% 123
1.50 768 12,288 401 4.90% 33
2.00 1024 16,384 0 0.00% 0
what the buffer costs: padded compute is capacity, not actual load
cf=1.0 buffer holds 8,192 slots for 8,192 assignments -> 100% of the buffer is real work
cf=1.25 buffer holds 10,240 slots for 8,192 assignments -> 80% of the buffer is real work
cf=2.0 buffer holds 16,384 slots for 8,192 assignments -> 50% of the buffer is real work
what a dropped token actually experiences
processed: [1.5, 1.5, 3.0999999046325684]
dropped : [1.0, 2.0, 3.0] <- the layer becomes the identityReading the output
Max-to-mean load is only 1.90, and capacity factor 1.0 still drops 19.6%. This is the number people get wrong. A capacity factor of 1.0 provides seats for the average expert, so every expert above average overflows. Mild imbalance produces severe dropping.
Capacity factor 2.0 drops nothing here, and half the buffer is empty. The busiest expert took 971 of 1024 slots; the quietest took 127. You pay for 1024 slots on all sixteen experts either way, because the buffer is uniform.
318 tokens got no expert at all at cf=1.0. The tokens with no expert column is the one that matters for quality. Losing one of two experts degrades a token; losing both makes the layer an identity function for it.
A dropped token comes out unchanged, not zeroed. [1.0, 2.0, 3.0] in and out. Because of the residual connection, dropping is silent. There is no exception, no NaN, no shape error. You will not find this without instrumenting for it.
The formula
capacity_per_expert = ceil( tokens_in_batch * top_k / num_experts * capacity_factor )Note top_k in the numerator: with top-2 routing there are twice as many assignments as tokens. Omitting it is a common bug that halves your buffers and doubles your drop rate.
Conventional settings, from GShard onwards: 1.25 during training, 2.0 during evaluation. Evaluation batches are smaller, so the law of large numbers helps less and imbalance is worse per batch.
Dropless training
MegaBlocks (Gale, Narayanan, Young and Zaharia, 2022) reformulates the MoE layer as block-sparse matrix multiplication, so ragged per-expert groups need no padding and no truncation. Their claim is direct: the system "never drops tokens and maps efficiently to modern hardware", with end-to-end training speed-ups of up to 40% over Tutel and 2.4x over Megatron-LM.
Today most large MoE training runs are dropless, via MegaBlocks-style block-sparse kernels or grouped GEMM implementations that take a list of per-expert row counts. If your framework offers it, use it: the capacity dial is a workaround for a kernel limitation that no longer exists.
Capacity survives at inference under expert parallelism, where fixed-size buffers make the all-to-all communication plan static and predictable. There, the trade in the table above is still live.
Common mistakes
Not logging the drop rate. The single most important instrument for an MoE run, and the easiest to omit because nothing fails.
Forgetting top_k in the capacity formula. Silently halves the buffer.
Using the training capacity factor at inference. Inference batches are smaller and less balanced. Raise the factor or use dropless kernels.
Assuming first-come-first-served ordering is neutral. With position-ordered dispatch, later tokens in the batch are dropped preferentially, so the ends of long sequences suffer more. Some implementations dispatch by routing probability instead, dropping the least-confident assignments — a better policy and not the default everywhere.
Comparing MoE quality across frameworks without checking the capacity setting. A model evaluated at cf=1.0 and the same model at cf=2.0 are different functions.
Try it yourself
Change torch.randn(E) * 0.6 to * 1.5 for a sharper router. Max-to-mean rises and the drop rate at every capacity factor rises with it. Then set K = 1 and confirm the tokens with no expert column now equals the dropped-slot count, because there is no second chance.
What to learn next
- Shared and fine-grained experts — smaller experts, and what that does to load statistics.
- Serving a MoE across GPUs — where fixed buffers still earn their keep.
- Load balancing and expert collapse — the upstream cause of every drop.
Researcher — Mathematics and papers.
Definition
Expert capacity per expert per batch:
$$ C = \left\lceil \frac{T \cdot k}{N} \cdot c \right\rceil $$
$T$ is tokens per batch (per expert-parallel group), $k$ the top-$k$, $N$ the expert count and $c$ the capacity factor. The dispatch tensor has fixed shape $[N, C]$, so the expert computation is a dense batched matmul of static shape — which is why the constraint exists at all.
Introduced by Lepikhin et al. (2020), GShard (arXiv:2006.16668) alongside top-2 routing and automatic sharding. Fedus et al. (2021), Switch Transformer studied $c$ directly and found top-1 routing with modest capacity factors to be a better trade than top-2 with tight capacity, at equal compute.
Why mild imbalance produces severe dropping
Let $n_i$ be the load on expert $i$ with $\sum_i n_i = Tk$. Dropped assignments are
$$ D = \sum_{i=1}^{N} \max(0,\; n_i - C) $$
This depends on the upper tail of the load distribution, not its variance. Even under a balanced router, $n_i$ fluctuates; with $\text{Var}(n_i) = \sigma^2$ and $C = \mu c$, drops appear whenever $n_i > \mu c$, which for $c = 1$ happens for roughly half of experts.
The experiment above quantifies it: max-to-mean of 1.90 gives 19.6% drops at $c = 1$. This is why $c = 1$ is never used, and why the balancing losses of the previous lesson exist as much for capacity as for parameter utilisation.
Gradient consequences
A dropped assignment contributes nothing to the output, so it contributes nothing to the gradient of the corresponding expert or of the router entry that selected it. Overloaded experts therefore receive less gradient signal per selection than their popularity implies — a mild self-correcting pressure, and far too weak to rely on.
Dropping also creates a train-test asymmetry. Batch composition determines which tokens are dropped, so the same token in a different batch takes a different path. The layer is a function of the batch, not of the token. Any evaluation with a different batch size is measuring a different function.
Dropless formulations
MegaBlocks (Gale, Narayanan, Young and Zaharia, 2022, arXiv:2211.15841) expresses the MoE layer as block-sparse matrix products with a variable-size block structure, so the per-expert group sizes need not be equal or padded. They report never dropping tokens with up to 40% end-to-end speed-up over Tutel and 2.4x over Megatron-LM.
The equivalent modern primitive is grouped GEMM — a single kernel taking a list of per-group row counts, available in CUTLASS, Triton and vendor libraries. Practically all recent large MoE training uses one of these paths, and capacity factor is absent from the configs of models like DeepSeek-V3.
Where capacity persists
Under expert parallelism, tokens are shipped between devices by all-to-all collectives. Variable-size all-to-all is supported but complicates buffer allocation, overlap scheduling and static graph capture. Fixed capacity makes the communication volume known in advance, which is worth real throughput.
The current split in practice:
| Setting | Typical approach |
|---|---|
| Large-scale pre-training | Dropless block-sparse or grouped GEMM |
| Single-node fine-tuning | Either; capacity is simpler |
| Production serving with expert parallelism | Fixed capacity, plus expert replication for hot experts |
Serving-side, the more effective lever is not the capacity factor but expert placement. DeepSeek's EPLB replicates hot experts across devices using measured load statistics, which reduces the tail that capacity has to absorb. Raising $c$ treats the symptom; replication treats the cause. See serving a MoE across GPUs.
Instrumentation
Log per layer, per step: drop rate, fraction of tokens receiving zero experts, max-to-mean load, and buffer utilisation. The second of these correlates with quality far better than the first, and is the one almost nobody records.
What to learn next
- Shared and fine-grained experts — smaller experts, and what that does to load statistics.
- Serving a MoE across GPUs — where fixed buffers still earn their keep.
- Load balancing and expert collapse — the upstream cause of every drop.