Mixture of Experts

Shared and fine-grained experts

Cutting experts into many smaller ones gives the router vastly more combinations at the same cost per token, and keeping one expert always-on saves every other expert from relearning the basics.

On this page 8
  1. The short answer
  2. The spice box in every kitchen
  3. Why that matters here
  4. The salt that never goes in a compartment
  5. What it costs
  6. The direction of travel
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Use many small experts instead of a few big ones, and keep one expert switched on for every word.

The spice box in every kitchen

A masala dabba with four big compartments lets you make a handful of dishes. Take two spices at a time and there are only a few pairs to choose from.

Cut the same total spice into sixteen small compartments and take eight of them. You use the same amount of spice, and the number of possible mixtures is now enormous.

Nothing was added. The same quantity, divided finer, gives far more recipes.

Why that matters here

An expert is a compartment. The router takes a few per word.

With eight experts and two picks there are twenty-eight possible pairs. With five hundred and twelve experts and ten picks, the count is astronomical. More combinations than grains of sand on a beach.

Each word can be handled by a mixture assembled precisely for it. The work per word has not changed at all.

The salt that never goes in a compartment

Now the second idea, and it comes from the same kitchen.

Every dish needs salt. You would not put salt into one numbered compartment and hope the cook picks that compartment today.

If you did, every compartment would end up holding a little salt, wasting space that could hold something distinctive.

So salt sits outside the box, always within reach. In a model this is called a shared expert. It runs for every single word, and holds the general knowledge all words need.

The routed experts are then free to be genuinely different from each other. None of them has to carry the basics.

What it costs

Small experts mean small pieces of work. Computers are less efficient at many small pieces than at a few big ones.

And more experts means the router has more scores to produce, and more places for a word to be sent. When experts live on different machines, that means more network traffic.

So there is a sweet spot. It has moved upwards every year, as the software gets better at handling many small experts.

The direction of travel

You can watch this happen across real models. An early open model used eight experts. A later one used thirty-two. Then a hundred and twenty-eight, then two hundred and fifty-six, then five hundred and twelve.

Almost all the recent ones also keep one shared expert switched on. The pattern is now close to standard.

Remember this

  • Many small experts give hugely more combinations at the same work per word.
  • A shared expert runs for every word and holds the common knowledge.
  • The limits are computer efficiency and network traffic, not the idea itself.

What to learn next

Developer — Code and libraries.

Setup

bash
python3 --version     # this one uses the standard library only

Written against Python 3.10. No packages needed.

Counting the combinations

granularity.py
from math import comb

d_model = 4096
base_ff = 14336                    # one ordinary feed-forward block

print("splitting experts finer at constant compute per token")
print(f"{'m':>3}{'experts':>9}{'top-k':>7}{'expert ff dim':>15}"
      f"{'active ff dim':>15}{'combinations':>32}")
N0, K0 = 16, 2
for m in (1, 2, 4, 8):
    N, K = N0 * m, K0 * m
    ff = base_ff // m
    print(f"{m:>3}{N:>9}{K:>7}{ff:>15,}{K*ff:>15,}{comb(N, K):>32,}")

print("\nsame idea at the scale real models use")
for name, N, K in [("Mixtral-8x7B", 8, 2), ("gpt-oss-20b", 32, 4),
                   ("Qwen3-30B-A3B", 128, 8), ("DeepSeek-V3", 256, 8),
                   ("Qwen3-Next-80B", 512, 10)]:
    print(f"  {name:<16} {N:>4} experts, top-{K:<3} -> {comb(N, K):,} possible combinations")

# A shared expert runs for every token. It is capacity you never have to route to.
print("\nDeepSeek-V3 feed-forward accounting, per MoE layer (from config.json)")
h, moe_ff, n_routed, n_shared, topk = 7168, 2048, 256, 1, 8
per_expert = 3 * h * moe_ff        # gate, up and down projections (SwiGLU)
total = (n_routed + n_shared) * per_expert + h * n_routed
active = (topk + n_shared) * per_expert + h * n_routed
print(f"  parameters per expert        : {per_expert:>15,}")
print(f"  experts held in memory       : {n_routed} routed + {n_shared} shared")
print(f"  total parameters in the layer: {total:>15,}")
print(f"  active per token             : {active:>15,}   ({active/total:.2%})")
print(f"  of which the shared expert   : {n_shared*per_expert:>15,}   "
      f"({n_shared/(topk+n_shared):.0%} of the active feed-forward work)")
Output
splitting experts finer at constant compute per token
  m  experts  top-k  expert ff dim  active ff dim                    combinations
  1       16      2         14,336         28,672                             120
  2       32      4          7,168         28,672                          35,960
  4       64      8          3,584         28,672                   4,426,165,368
  8      128     16          1,792         28,672      93,343,021,201,262,177,400

same idea at the scale real models use
  Mixtral-8x7B        8 experts, top-2   -> 28 possible combinations
  gpt-oss-20b        32 experts, top-4   -> 35,960 possible combinations
  Qwen3-30B-A3B     128 experts, top-8   -> 1,429,702,652,400 possible combinations
  DeepSeek-V3       256 experts, top-8   -> 409,663,695,276,000 possible combinations
  Qwen3-Next-80B    512 experts, top-10  -> 312,268,282,598,377,321,216 possible combinations

DeepSeek-V3 feed-forward accounting, per MoE layer (from config.json)
  parameters per expert        :      44,040,192
  experts held in memory       : 256 routed + 1 shared
  total parameters in the layer:  11,320,164,352
  active per token             :     398,196,736   (3.52%)
  of which the shared expert   :      44,040,192   (11% of the active feed-forward work)

Reading the output

The active ff dim column never changes: 28,672 at every granularity. That is the constraint the whole idea is built around. Splitting experts by m and multiplying top_k by m holds compute per token exactly fixed.

Combinations go from 120 to 93 quintillion. The router's expressiveness — how many distinct feed-forward functions it can assemble — grows super-exponentially while the compute bill stays flat. These are exactly the numbers in the DeepSeekMoE paper, which uses the 16-expert-top-2 to 64-expert-top-8 comparison.

Mixtral has 28 combinations. Twenty-eight distinct pairs, for the entire English language, at every layer. Seen that way, the move to hundreds of experts looks less like an optimisation and more like a correction.

DeepSeek-V3 activates 3.52% of each MoE layer. 398 million parameters of 11.3 billion, per token, per layer.

The shared expert is 11% of the active feed-forward work. One of nine active experts. It is small in compute and structurally important, because it is the only expert guaranteed to see every token, so it is the only one that can learn from the whole distribution.

The granularity dial in real configs

Modelexpertstop-kexpert FF dimshared
Mixtral-8x7B-v0.182143360
gpt-oss-20b32428800
Qwen3-30B-A3B12887680
DeepSeek-V3256820481
Qwen3-Next-80B-A3B512105121

Read the expert FF dim column against the dense equivalent. Mixtral's experts are full-size feed-forward blocks. Qwen3-Next's are 512 wide — a twenty-eighth of Mixtral's, with 64 times as many of them.

Note also first_k_dense_replace: 3 in DeepSeek-V3: the first three layers keep an ordinary dense feed-forward block of width 18432. Routing in the earliest layers is unstable, and several families now do this.

Why not go finer still

Matrix multiplies get small. A 512-wide expert processing a handful of tokens is a tiny GEMM. GPUs reach peak throughput on large ones. Beyond some granularity you are memory-bound on weight loading rather than compute-bound — see memory-bound vs compute-bound.

The router grows. Router parameters are d_model x num_experts per layer, and its softmax and top-k run over every expert. At 512 experts this is still small, and it is no longer free.

All-to-all traffic grows. With top-10 over 512 experts spread across nodes, a token's activations may need to reach ten different machines. This is why DeepSeek adds node-limited routing, capping the machines any one token touches.

Load statistics get noisier. With 512 experts and a modest batch, per-expert counts are small and their relative variance is large, which worsens both balancing and capacity behaviour.

Common mistakes

Increasing the expert count without increasing top-k. That reduces compute per token as well as expressiveness, so the comparison to the previous model is not like-for-like. Granularity means scaling both together.

Assuming the shared expert is optional plumbing. It changes what the routed experts learn. Removing it and continuing training does not give you back the same model.

Sizing memory from the active count. All 257 experts per layer are resident. Fine-grained models are more memory-hungry per unit of active compute, not less.

Comparing expert counts across models without expert width. 512 experts of width 512 and 8 experts of width 14336 are different designs, not different points on one axis.

Try it yourself

Add m = 16 to the loop and watch the expert FF dim fall to 896. Then compute the router parameter count d_model * N at each granularity and find where it stops being a rounding error against the expert parameters.

What to learn next

Researcher — Mathematics and papers.

The two strategies

Dai et al. (2024), DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models (arXiv:2401.06066) state both precisely.

Fine-grained expert segmentation: "finely segmenting the experts into $mN$ ones and activating $mK$ from them", which preserves both parameter count and computation per token while multiplying the number of achievable expert combinations from $\binom{N}{K}$ to $\binom{mN}{mK}$.

Shared expert isolation: "isolating $K_s$ experts as shared ones, aiming at capturing common knowledge and mitigating redundancy in routed experts."

Their reported results: DeepSeekMoE 2B matches GShard 2.9B; DeepSeekMoE 16B matches LLaMA2 7B at about 40% of the computation; DeepSeekMoE 145B approaches DeepSeek 67B at a comparable fraction.

Why segmentation helps

The MoE layer computes $\sum_{i \in \mathcal{S}} g_i E_i(x)$ over the selected set $\mathcal{S}$. The reachable function class is indexed by $\mathcal{S}$, so its size is $\binom{N}{K}$.

Under segmentation with factor $m$, each expert's hidden width scales as $1/m$ and $K \to mK$. Active parameters $mK \cdot d_{ff}/m = K d_{ff}$ are invariant, while

$$ \binom{mN}{mK} \gg \binom{N}{K} $$

grows super-exponentially in $m$. The intuition is that a coarse expert must be a jack of all trades because it is selected as an indivisible unit, whereas fine experts can be composed.

The counter-pressure is statistical, not combinatorial: each fine expert sees roughly the same number of tokens per batch (since $mK$ grows with $mN$), but has $1/m$ the parameters, so per-parameter gradient signal is unchanged while per-expert load variance rises.

Why shared experts help

Without a shared expert, every routed expert must independently learn features common to all tokens — basic syntax, frequency statistics, the identity-like components of the residual update. That duplication consumes capacity in all $N$ experts.

Isolating $K_s$ always-on experts factorises the layer into a dense component and a sparse one:

$$ y = \underbrace{\sum_{s=1}^{K_s} E_s(x)}{\text{always on}} + \underbrace{\sum{i \in \operatorname{Top-}K} g_i E_i(x)}_{\text{routed}} $$

The dense component absorbs the shared structure and the routed component becomes a residual specialised to the token. This also stabilises training: the layer has a guaranteed non-zero output even if routing is degenerate or every routed assignment is dropped for capacity.

Current practice is $K_s = 1$ almost universally. Llama 4 takes it further at the other extreme, pairing one shared expert with routing to a single routed expert.

Scaling laws for granularity

Krajewski et al. (2024), Scaling Laws for Fine-Grained Mixture of Experts (arXiv:2402.07871) introduce granularity as an explicit hyperparameter — "whose adjustment enables precise control over the size of the experts" — and fit scaling laws over model size, training tokens and granularity.

Two conclusions worth carrying:

"The common practice of setting the size of experts in MoE to mirror the feed-forward layer is not optimal at almost any computational budget." Mixtral's design, expert width equal to the dense FFN width, is the practice this refutes.

"The efficiency gap between dense and MoE models widens as we scale up the model size and training budget." MoE is not a fixed-percentage saving; its advantage grows with scale, which is why it dominates at the frontier and is often unremarkable at small scale.

Systems limits

Granularity is bounded by hardware, not by the scaling law.

  • GEMM efficiency. Expert widths near 512 with small per-expert token counts underutilise tensor cores. Grouped GEMM kernels mitigate but do not remove this.
  • All-to-all volume. Under expert parallelism, communication scales with $K$ and with the number of distinct devices touched. Node-limited routing (DeepSeek-V3's n_group: 8, topk_group: 4) bounds the second term.
  • Load variance. Per-expert token counts shrink as $N$ grows at fixed batch, so relative variance grows as $1/\sqrt{n_i}$, degrading balance and raising drop rates at fixed capacity factor.

The observed progression — 8, 32, 128, 256, 512 experts across successive model generations — tracks improvements in these kernels and collectives more closely than it tracks any new modelling insight.

What to learn next