AI glossary

Mixture of experts (MoE)

In one sentence A mixture-of-experts model holds many specialist sub-networks but activates only a few per token, buying huge capacity at a fraction of the compute.

By Updated

A mixture-of-experts model contains many parallel sub-networks ("experts") plus a router that sends each token to only a few of them — so most of the model stays idle on any given token.

A large hospital does not send every patient through every department. Reception looks at you and routes: chest pain to cardiology, a rash to dermatology. The hospital's total expertise is enormous; the effort spent per patient involves two or three departments. An MoE model is that hospital. A dense model, by contrast, is a clinic where every patient must see every doctor.

Inside a transformer, the swap happens in the feed-forward layers: instead of one big feed-forward block, an MoE layer holds, say, 8 to 256 of them, and a small learned router picks the top 1 or 2 per token. The result decouples two things that dense models tie together: total parameters (knowledge capacity) and compute per token (cost).

Mixtral 8x7B : 47B total parameters, ~13B active per token
DeepSeek-V3  : 671B total, ~37B active per token

The catches are operational. All parameters must sit in memory even though few run, so MoE trades cheap compute for expensive VRAM. Training needs load-balancing tricks so the router does not overuse a few favourite experts. And the "experts" are not human-readable specialists — the router's division of labour is statistical, not thematic. GPT-4 is widely reported to be MoE; Mixtral and DeepSeek made the recipe open.

Where to go next