Mixture of Experts

Active parameters vs total parameters

A mixture-of-experts model has two sizes, one that sets your memory bill and one that sets your speed, and quoting either alone gives a misleading picture.

On this page 8
  1. The short answer
  2. The library and the reading
  3. What each number decides
  4. Reading the names
  5. The comparison people get wrong
  6. Who this suits and who it does not
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A mixture-of-experts model has two sizes. One is how much you must store. The other is how much runs for each word.

The library and the reading

A library has forty thousand books. To open the library you need the building, the shelves and every book on them.

To answer one question you read three books. Your afternoon is set by the three. The rent is set by the forty thousand.

Nobody would describe that library with one number. Models like these need two numbers for the same reason.

What each number decides

Total parameters is the whole library. It decides how much memory you need, how many machines, and how much the hardware costs.

Active parameters is the three books. It decides how fast each word comes out, and how much electricity each word burns.

   DeepSeek-V3

   total  : 671 billion   ->  about 1250 GB to hold in memory
   active :  37 billion   ->  the speed of a much smaller model

   you rent the hardware of a 671 billion model
   you get the speed of a 37 billion model

Reading the names

Model names now often carry both numbers. A name ending in "80B-A3B" means eighty billion in total, three billion active.

Older names hid it. A name like "8x7B" suggests fifty-six billion. The real total is about forty-seven billion, because the parts outside the experts are shared rather than copied eight times.

Multiplying the two numbers in that style of name gives you the wrong answer. It is worth knowing before you plan a purchase.

The comparison people get wrong

Someone benchmarks a big sparse model against a small dense one. They come out similar in speed, so the big one looks disappointing.

Or they compare it to a dense model of the same total size. It comes out worse in quality, so the same conclusion follows.

Both comparisons are wrong. The honest position is that this design sits between the two. Better quality than the small dense model, far cheaper to run than the large one, and costly to house.

Who this suits and who it does not

A busy shared service holds the model once and serves thousands of people. The memory cost is paid once and spread thin. The speed benefit lands on every request.

One person with one graphics card gets the opposite deal. The memory cost is the whole problem and the speed benefit helps one user.

That is why these models dominate hosted services and are awkward to run at home.

Remember this

  • Total sets memory and hardware. Active sets speed and energy per word.
  • The two are quoted together in modern names, and the ratio is often ten to twenty times.
  • The design suits big shared servers far better than a single machine.

What to learn next

Developer — Code and libraries.

Setup

bash
python3 --version     # standard library only

Python 3.10. No packages.

Rebuilding a published parameter count from the config

Every number below comes from deepseek-ai/DeepSeek-V3's config.json. Nothing is fitted or estimated.

param_count.py
# Every number below is read from deepseek-ai/DeepSeek-V3's config.json.
L, dense_layers      = 61, 3
h, vocab             = 7168, 129280
dense_ff, moe_ff     = 18432, 2048
n_routed, n_shared, topk = 256, 1, 8
heads                = 128
q_lora, kv_lora      = 1536, 512
qk_nope, qk_rope, v_dim = 128, 64, 128

def mla_params():
    return (h * q_lora                       # down-project the query
            + q_lora * heads * (qk_nope + qk_rope)   # up-project it
            + h * kv_lora                    # down-project keys and values
            + h * qk_rope                    # the shared rotary key
            + kv_lora * heads * qk_nope      # up-project keys
            + kv_lora * heads * v_dim        # up-project values
            + heads * v_dim * h)             # output projection

expert   = 3 * h * moe_ff                    # SwiGLU: gate, up, down
dense_fn = 3 * h * dense_ff
router   = h * n_routed

attn   = L * mla_params()
embeds = 2 * vocab * h                       # input embedding and output head, untied
ffn_total  = dense_layers * dense_fn + (L - dense_layers) * ((n_routed + n_shared) * expert + router)
ffn_active = dense_layers * dense_fn + (L - dense_layers) * ((topk + n_shared) * expert + router)

total  = attn + embeds + ffn_total
active = attn + vocab * h + ffn_active       # one embedding row is looked up, not multiplied

print(f"{'component':<32}{'total':>18}{'active per token':>20}")
print(f"{'attention (MLA), 61 layers':<32}{attn:>18,}{attn:>20,}")
print(f"{'embedding + output head':<32}{embeds:>18,}{vocab*h:>20,}")
print(f"{'feed-forward (3 dense + 58 MoE)':<32}{ffn_total:>18,}{ffn_active:>20,}")
print(f"{'-'*70}")
print(f"{'TOTAL':<32}{total:>18,}{active:>20,}")
print(f"\n  {total/1e9:.0f}B total, {active/1e9:.0f}B active   "
      f"-> {active/total:.1%} of the model runs per token")
print("  the published figures are 671B total and 37B activated per token")

print("\nwhat each number is for")
print(f"  memory to hold the weights in bf16 : {total*2/2**30:>8.0f} GiB")
print(f"  arithmetic per token (2 * active)  : {2*active/1e9:>8.0f} GFLOPs")
print(f"  a dense model with the same FLOPs  : {active/1e9:>8.0f}B parameters")
print(f"  a dense model with the same memory : {total/1e9:>8.0f}B parameters")

print("\nnaming conventions you will meet")
for nm, tot, act in [("Mixtral-8x7B", 46.7, 12.9), ("gpt-oss-120b", 116.8, 5.1),
                     ("gpt-oss-20b", 21.0, 3.6), ("Qwen3-30B-A3B", 30.0, 3.0),
                     ("Qwen3-Next-80B-A3B", 80.0, 3.0), ("DeepSeek-V3", 671.0, 37.0)]:
    print(f"  {nm:<20}{tot:>7.1f}B total{act:>7.1f}B active   ratio {tot/act:>5.1f}x")
Output
component                                    total    active per token
attention (MLA), 61 layers          11,413,422,080      11,413,422,080
embedding + output head              1,853,358,080         926,679,040
feed-forward (3 dense + 58 MoE)    657,758,617,600      24,284,495,872
----------------------------------------------------------------------
TOTAL                              671,025,397,760      36,624,596,992

  671B total, 37B active   -> 5.5% of the model runs per token
  the published figures are 671B total and 37B activated per token

what each number is for
  memory to hold the weights in bf16 :     1250 GiB
  arithmetic per token (2 * active)  :       73 GFLOPs
  a dense model with the same FLOPs  :       37B parameters
  a dense model with the same memory :      671B parameters

naming conventions you will meet
  Mixtral-8x7B           46.7B total   12.9B active   ratio   3.6x
  gpt-oss-120b          116.8B total    5.1B active   ratio  22.9x
  gpt-oss-20b            21.0B total    3.6B active   ratio   5.8x
  Qwen3-30B-A3B          30.0B total    3.0B active   ratio  10.0x
  Qwen3-Next-80B-A3B     80.0B total    3.0B active   ratio  26.7x
  DeepSeek-V3           671.0B total   37.0B active   ratio  18.1x

Reading the output

671,025,397,760 against a published 671B, and 36.6B against a published 37B. The config alone reproduces both headline numbers. That is worth doing yourself for any model you are about to deploy: it catches misreported figures and it forces you to understand where the parameters actually live.

The feed-forward row carries 98% of the total and 66% of the active. Attention is 11.4B and is fully dense — every parameter runs for every token. That is why the active fraction is 5.5% rather than the 3.5% you would get from the MoE layers alone.

1250 GiB of weights. At bf16 that needs sixteen 80 GB GPUs before a single token of KV cache. The model is deployed at reduced precision in practice, and it remains a multi-node model.

73 GFLOPs per token. Roughly what a 37B dense model costs. That is the trade in one line: 1250 GiB of memory buying 37B-model speed with, in benchmarks, far better than 37B-model quality.

gpt-oss-120b has a 22.9x ratio; Mixtral has 3.6x. Two years apart, and the sparsity has grown by a factor of six. That progression is the single clearest trend in this section.

The rule for estimating

memory (bytes)        ~ total_params * bytes_per_param     + KV cache + activations
compute per token     ~ 2 * active_params  FLOPs
decode time per token ~ (active_weight_bytes + kv_bytes) / memory_bandwidth

The third line matters more than the second. Decoding is memory-bound (see memory-bound vs compute-bound), so decode speed depends on the bytes of active weights read per token, not on FLOPs.

There is a wrinkle that trips up capacity planning. At batch size 1 you read only the active experts. At batch size 256, different tokens select different experts, and in the limit you read every expert every step. An MoE model's per-token cost therefore rises with batch size in a way a dense model's does not, until it saturates at the full weight read.

Common mistakes

Multiplying out "8x7B". Mixtral-8x7B is about 46.7B, not 56B. Attention, embeddings and norms are shared across experts rather than replicated.

Sizing GPUs from the active count. The most expensive mistake in this lesson. Every expert is resident.

Comparing an MoE model to a dense model of equal total size. DeepSeek-V3 is not competing with a 671B dense model. Nobody has trained one.

Using training FLOPs rules of thumb unchanged. 6ND uses active parameters for N, so an MoE model trains far more cheaply per token than its total size suggests. That is exactly why the design is used.

Ignoring the router and shared expert in the active count. Small terms, but the shared expert is a ninth of DeepSeek-V3's active feed-forward work.

Try it yourself

Change topk from 8 to 4 and recompute. Active parameters fall while total is unchanged, so the ratio doubles. Then set n_shared = 0 and see how much of the active count the shared expert was carrying.

What to learn next

Researcher — Mathematics and papers.

Definitions

$$ P_{\text{total}} = P_{\text{dense}} + L_{\text{moe}} \left(N + K_s\right) P_e, \qquad P_{\text{active}} = P_{\text{dense}} + L_{\text{moe}} \left(k + K_s\right) P_e + L_{\text{moe}} d N $$

$P_{\text{dense}}$ covers attention, embeddings, norms and any dense feed-forward layers. $N$ is routed experts, $K_s$ shared experts, $k$ the top-$k$, $P_e$ parameters per expert and $dN$ the router. The sparsity ratio $P_{\text{total}}/P_{\text{active}}$ is the design's headline knob.

What each quantity governs

QuantityGoverned byPractical consequence
Weight memory$P_{\text{total}}$GPU count, cost floor
Training FLOPs$\approx 6 P_{\text{active}} D$Cost of a pre-training run
Prefill FLOPs$\approx 2 P_{\text{active}} S$Time to first token
Decode bandwidthactive weight bytes readTokens per second per stream
Aggregate throughputactive weight bytes and all-to-all volumeCost per served token

Decode is the subtle one. At batch $B$, the set of experts touched is the union over $B$ tokens of their top-$k$ selections. For $N$ experts under near-uniform routing, the expected number of distinct experts touched is

$$ \mathbb{E}[|\mathcal{U}|] = N\left(1 - \left(1 - \tfrac{k}{N}\right)^{B}\right) $$

which approaches $N$ quickly. At $N = 256$, $k = 8$ and $B = 128$, the expectation is already above 250. Beyond modest batch sizes, an MoE model reads essentially all its weights per step and the active-parameter saving in bandwidth terms is gone; what remains is the FLOP saving and the fact that each expert's read is amortised over more tokens.

This is why MoE serving economics depend so heavily on batch size and expert placement, and why "active parameters" is a training-cost and low-batch-latency statistic rather than a universal one.

Scaling laws

Clark et al. (2022), Unified Scaling Laws for Routed Language Models fit loss as a function of dense-equivalent parameters and expert count, finding gains that saturate with $N$ at the scales studied.

Krajewski et al. (2024), Scaling Laws for Fine-Grained Mixture of Experts (arXiv:2402.07871) add granularity and reach a stronger conclusion: "MoE models consistently outperform dense Transformers", and "the efficiency gap between dense and MoE models widens as we scale up the model size and training budget."

That widening gap is the reason every frontier open model released since 2024 is sparse. It is not a fixed-percentage efficiency; it compounds with scale.

Reporting standards

A defensible model card gives all of:

  • total parameters, active parameters per token, and the expert configuration ($N$, $k$, $K_s$, expert width)
  • which layers are dense (first_k_dense_replace or equivalent)
  • whether attention is dense, GQA or latent, with its own parameter count
  • measured throughput at stated batch sizes, precision and hardware

The gap between an MoE model's advertised active count and its realised serving cost is the largest and least-documented number in current model releases. If it is not in the card, compute it: the config file is sufficient, as the script above demonstrates.

The name of the trade

An MoE model converts a memory budget into a quality improvement at fixed compute. That is only a good trade where memory is abundant relative to compute demand — large shared serving fleets, and training clusters with high aggregate memory.

Where memory is the binding constraint, the trade inverts. A single consumer GPU running Qwen3-30B-A3B holds 30B of weights to obtain 3B of compute per token, and the 27B of resident-but-idle parameters are pure cost. The right comparison there is a dense model of the same memory footprint, and the sparse model does not win it on any evidence I know of.

What to learn next