Mixture of experts (MoE)
In one sentence A mixture-of-experts model holds many specialist sub-networks but activates only a few per token, buying huge capacity at a fraction of the compute.
Updated
A mixture-of-experts model contains many parallel sub-networks ("experts") plus a router that sends each token to only a few of them — so most of the model stays idle on any given token.
A large hospital does not send every patient through every department. Reception looks at you and routes: chest pain to cardiology, a rash to dermatology. The hospital's total expertise is enormous; the effort spent per patient involves two or three departments. An MoE model is that hospital. A dense model, by contrast, is a clinic where every patient must see every doctor.
Inside a transformer, the swap happens in the feed-forward layers: instead of one big feed-forward block, an MoE layer holds, say, 8 to 256 of them, and a small learned router picks the top 1 or 2 per token. The result decouples two things that dense models tie together: total parameters (knowledge capacity) and compute per token (cost).
Mixtral 8x7B : 47B total parameters, ~13B active per token
DeepSeek-V3 : 671B total, ~37B active per tokenThe catches are operational. All parameters must sit in memory even though few run, so MoE trades cheap compute for expensive VRAM. Training needs load-balancing tricks so the router does not overuse a few favourite experts. And the "experts" are not human-readable specialists — the router's division of labour is statistical, not thematic. GPT-4 is widely reported to be MoE; Mixtral and DeepSeek made the recipe open.
Where to go next
- Full lesson: How LLMs work
- Related terms: transformer, parameter, inference, foundation-model