Active parameters vs total parameters
A mixture-of-experts model has two sizes, one that sets your memory bill and one that sets your speed, and quoting either alone gives a misleading picture.
- 11 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A mixture-of-experts model has two sizes. One is how much you must store. The other is how much runs for each word.
The library and the reading
A library has forty thousand books. To open the library you need the building, the shelves and every book on them.
To answer one question you read three books. Your afternoon is set by the three. The rent is set by the forty thousand.
Nobody would describe that library with one number. Models like these need two numbers for the same reason.
What each number decides
Total parameters is the whole library. It decides how much memory you need, how many machines, and how much the hardware costs.
Active parameters is the three books. It decides how fast each word comes out, and how much electricity each word burns.
DeepSeek-V3
total : 671 billion -> about 1250 GB to hold in memory
active : 37 billion -> the speed of a much smaller model
you rent the hardware of a 671 billion model
you get the speed of a 37 billion modelReading the names
Model names now often carry both numbers. A name ending in "80B-A3B" means eighty billion in total, three billion active.
Older names hid it. A name like "8x7B" suggests fifty-six billion. The real total is about forty-seven billion, because the parts outside the experts are shared rather than copied eight times.
Multiplying the two numbers in that style of name gives you the wrong answer. It is worth knowing before you plan a purchase.
The comparison people get wrong
Someone benchmarks a big sparse model against a small dense one. They come out similar in speed, so the big one looks disappointing.
Or they compare it to a dense model of the same total size. It comes out worse in quality, so the same conclusion follows.
Both comparisons are wrong. The honest position is that this design sits between the two. Better quality than the small dense model, far cheaper to run than the large one, and costly to house.
Who this suits and who it does not
A busy shared service holds the model once and serves thousands of people. The memory cost is paid once and spread thin. The speed benefit lands on every request.
One person with one graphics card gets the opposite deal. The memory cost is the whole problem and the speed benefit helps one user.
That is why these models dominate hosted services and are awkward to run at home.
Remember this
- Total sets memory and hardware. Active sets speed and energy per word.
- The two are quoted together in modern names, and the ratio is often ten to twenty times.
- The design suits big shared servers far better than a single machine.
What to learn next
- Serving a MoE across GPUs — turning the total-parameter bill into a deployment plan.
- Quantization in practice — the standard way to cut the memory half of this trade.
- Latency and throughput — the two metrics active and total parameters map onto.
Developer — Code and libraries.
Setup
python3 --version # standard library onlyPython 3.10. No packages.
Rebuilding a published parameter count from the config
Every number below comes from deepseek-ai/DeepSeek-V3's config.json. Nothing is fitted or estimated.
# Every number below is read from deepseek-ai/DeepSeek-V3's config.json.
L, dense_layers = 61, 3
h, vocab = 7168, 129280
dense_ff, moe_ff = 18432, 2048
n_routed, n_shared, topk = 256, 1, 8
heads = 128
q_lora, kv_lora = 1536, 512
qk_nope, qk_rope, v_dim = 128, 64, 128
def mla_params():
return (h * q_lora # down-project the query
+ q_lora * heads * (qk_nope + qk_rope) # up-project it
+ h * kv_lora # down-project keys and values
+ h * qk_rope # the shared rotary key
+ kv_lora * heads * qk_nope # up-project keys
+ kv_lora * heads * v_dim # up-project values
+ heads * v_dim * h) # output projection
expert = 3 * h * moe_ff # SwiGLU: gate, up, down
dense_fn = 3 * h * dense_ff
router = h * n_routed
attn = L * mla_params()
embeds = 2 * vocab * h # input embedding and output head, untied
ffn_total = dense_layers * dense_fn + (L - dense_layers) * ((n_routed + n_shared) * expert + router)
ffn_active = dense_layers * dense_fn + (L - dense_layers) * ((topk + n_shared) * expert + router)
total = attn + embeds + ffn_total
active = attn + vocab * h + ffn_active # one embedding row is looked up, not multiplied
print(f"{'component':<32}{'total':>18}{'active per token':>20}")
print(f"{'attention (MLA), 61 layers':<32}{attn:>18,}{attn:>20,}")
print(f"{'embedding + output head':<32}{embeds:>18,}{vocab*h:>20,}")
print(f"{'feed-forward (3 dense + 58 MoE)':<32}{ffn_total:>18,}{ffn_active:>20,}")
print(f"{'-'*70}")
print(f"{'TOTAL':<32}{total:>18,}{active:>20,}")
print(f"\n {total/1e9:.0f}B total, {active/1e9:.0f}B active "
f"-> {active/total:.1%} of the model runs per token")
print(" the published figures are 671B total and 37B activated per token")
print("\nwhat each number is for")
print(f" memory to hold the weights in bf16 : {total*2/2**30:>8.0f} GiB")
print(f" arithmetic per token (2 * active) : {2*active/1e9:>8.0f} GFLOPs")
print(f" a dense model with the same FLOPs : {active/1e9:>8.0f}B parameters")
print(f" a dense model with the same memory : {total/1e9:>8.0f}B parameters")
print("\nnaming conventions you will meet")
for nm, tot, act in [("Mixtral-8x7B", 46.7, 12.9), ("gpt-oss-120b", 116.8, 5.1),
("gpt-oss-20b", 21.0, 3.6), ("Qwen3-30B-A3B", 30.0, 3.0),
("Qwen3-Next-80B-A3B", 80.0, 3.0), ("DeepSeek-V3", 671.0, 37.0)]:
print(f" {nm:<20}{tot:>7.1f}B total{act:>7.1f}B active ratio {tot/act:>5.1f}x")component total active per token attention (MLA), 61 layers 11,413,422,080 11,413,422,080 embedding + output head 1,853,358,080 926,679,040 feed-forward (3 dense + 58 MoE) 657,758,617,600 24,284,495,872 ---------------------------------------------------------------------- TOTAL 671,025,397,760 36,624,596,992 671B total, 37B active -> 5.5% of the model runs per token the published figures are 671B total and 37B activated per token what each number is for memory to hold the weights in bf16 : 1250 GiB arithmetic per token (2 * active) : 73 GFLOPs a dense model with the same FLOPs : 37B parameters a dense model with the same memory : 671B parameters naming conventions you will meet Mixtral-8x7B 46.7B total 12.9B active ratio 3.6x gpt-oss-120b 116.8B total 5.1B active ratio 22.9x gpt-oss-20b 21.0B total 3.6B active ratio 5.8x Qwen3-30B-A3B 30.0B total 3.0B active ratio 10.0x Qwen3-Next-80B-A3B 80.0B total 3.0B active ratio 26.7x DeepSeek-V3 671.0B total 37.0B active ratio 18.1x
Reading the output
671,025,397,760 against a published 671B, and 36.6B against a published 37B. The config alone reproduces both headline numbers. That is worth doing yourself for any model you are about to deploy: it catches misreported figures and it forces you to understand where the parameters actually live.
The feed-forward row carries 98% of the total and 66% of the active. Attention is 11.4B and is fully dense — every parameter runs for every token. That is why the active fraction is 5.5% rather than the 3.5% you would get from the MoE layers alone.
1250 GiB of weights. At bf16 that needs sixteen 80 GB GPUs before a single token of KV cache. The model is deployed at reduced precision in practice, and it remains a multi-node model.
73 GFLOPs per token. Roughly what a 37B dense model costs. That is the trade in one line: 1250 GiB of memory buying 37B-model speed with, in benchmarks, far better than 37B-model quality.
gpt-oss-120b has a 22.9x ratio; Mixtral has 3.6x. Two years apart, and the sparsity has grown by a factor of six. That progression is the single clearest trend in this section.
The rule for estimating
memory (bytes) ~ total_params * bytes_per_param + KV cache + activations
compute per token ~ 2 * active_params FLOPs
decode time per token ~ (active_weight_bytes + kv_bytes) / memory_bandwidthThe third line matters more than the second. Decoding is memory-bound (see memory-bound vs compute-bound), so decode speed depends on the bytes of active weights read per token, not on FLOPs.
There is a wrinkle that trips up capacity planning. At batch size 1 you read only the active experts. At batch size 256, different tokens select different experts, and in the limit you read every expert every step. An MoE model's per-token cost therefore rises with batch size in a way a dense model's does not, until it saturates at the full weight read.
Common mistakes
Multiplying out "8x7B". Mixtral-8x7B is about 46.7B, not 56B. Attention, embeddings and norms are shared across experts rather than replicated.
Sizing GPUs from the active count. The most expensive mistake in this lesson. Every expert is resident.
Comparing an MoE model to a dense model of equal total size. DeepSeek-V3 is not competing with a 671B dense model. Nobody has trained one.
Using training FLOPs rules of thumb unchanged. 6ND uses active parameters for N, so an MoE model trains far more cheaply per token than its total size suggests. That is exactly why the design is used.
Ignoring the router and shared expert in the active count. Small terms, but the shared expert is a ninth of DeepSeek-V3's active feed-forward work.
Try it yourself
Change topk from 8 to 4 and recompute. Active parameters fall while total is unchanged, so the ratio doubles. Then set n_shared = 0 and see how much of the active count the shared expert was carrying.
What to learn next
- Serving a MoE across GPUs — turning the total-parameter bill into a deployment plan.
- Quantization in practice — the standard way to cut the memory half of this trade.
- Latency and throughput — the two metrics active and total parameters map onto.
Researcher — Mathematics and papers.
Definitions
$$ P_{\text{total}} = P_{\text{dense}} + L_{\text{moe}} \left(N + K_s\right) P_e, \qquad P_{\text{active}} = P_{\text{dense}} + L_{\text{moe}} \left(k + K_s\right) P_e + L_{\text{moe}} d N $$
$P_{\text{dense}}$ covers attention, embeddings, norms and any dense feed-forward layers. $N$ is routed experts, $K_s$ shared experts, $k$ the top-$k$, $P_e$ parameters per expert and $dN$ the router. The sparsity ratio $P_{\text{total}}/P_{\text{active}}$ is the design's headline knob.
What each quantity governs
| Quantity | Governed by | Practical consequence |
|---|---|---|
| Weight memory | $P_{\text{total}}$ | GPU count, cost floor |
| Training FLOPs | $\approx 6 P_{\text{active}} D$ | Cost of a pre-training run |
| Prefill FLOPs | $\approx 2 P_{\text{active}} S$ | Time to first token |
| Decode bandwidth | active weight bytes read | Tokens per second per stream |
| Aggregate throughput | active weight bytes and all-to-all volume | Cost per served token |
Decode is the subtle one. At batch $B$, the set of experts touched is the union over $B$ tokens of their top-$k$ selections. For $N$ experts under near-uniform routing, the expected number of distinct experts touched is
$$ \mathbb{E}[|\mathcal{U}|] = N\left(1 - \left(1 - \tfrac{k}{N}\right)^{B}\right) $$
which approaches $N$ quickly. At $N = 256$, $k = 8$ and $B = 128$, the expectation is already above 250. Beyond modest batch sizes, an MoE model reads essentially all its weights per step and the active-parameter saving in bandwidth terms is gone; what remains is the FLOP saving and the fact that each expert's read is amortised over more tokens.
This is why MoE serving economics depend so heavily on batch size and expert placement, and why "active parameters" is a training-cost and low-batch-latency statistic rather than a universal one.
Scaling laws
Clark et al. (2022), Unified Scaling Laws for Routed Language Models fit loss as a function of dense-equivalent parameters and expert count, finding gains that saturate with $N$ at the scales studied.
Krajewski et al. (2024), Scaling Laws for Fine-Grained Mixture of Experts (arXiv:2402.07871) add granularity and reach a stronger conclusion: "MoE models consistently outperform dense Transformers", and "the efficiency gap between dense and MoE models widens as we scale up the model size and training budget."
That widening gap is the reason every frontier open model released since 2024 is sparse. It is not a fixed-percentage efficiency; it compounds with scale.
Reporting standards
A defensible model card gives all of:
- total parameters, active parameters per token, and the expert configuration ($N$, $k$, $K_s$, expert width)
- which layers are dense (
first_k_dense_replaceor equivalent) - whether attention is dense, GQA or latent, with its own parameter count
- measured throughput at stated batch sizes, precision and hardware
The gap between an MoE model's advertised active count and its realised serving cost is the largest and least-documented number in current model releases. If it is not in the card, compute it: the config file is sufficient, as the script above demonstrates.
The name of the trade
An MoE model converts a memory budget into a quality improvement at fixed compute. That is only a good trade where memory is abundant relative to compute demand — large shared serving fleets, and training clusters with high aggregate memory.
Where memory is the binding constraint, the trade inverts. A single consumer GPU running Qwen3-30B-A3B holds 30B of weights to obtain 3B of compute per token, and the 27B of resident-but-idle parameters are pure cost. The right comparison there is a dense model of the same memory footprint, and the sparse model does not win it on any evidence I know of.
What to learn next
- Serving a MoE across GPUs — turning the total-parameter bill into a deployment plan.
- Quantization in practice — the standard way to cut the memory half of this trade.
- Latency and throughput — the two metrics active and total parameters map onto.