Definition
Plain language
A model design where only a fraction of the parameters fire on any one input, letting models be very large but cheap to run.
As stated in the literature
A neural network architecture in which only a sparse subset of expert sub-networks is activated per token, enabling large total parameter counts at lower per-token compute.
Also called: MoE, mixture of experts, sparse mixture-of-experts
Why it matters: It lets models grow in total knowledge without proportionally growing the compute needed per token at inference.
For example, a 200-billion-parameter MoE model might only activate ~10 billion parameters when answering a single question.
Heard on the show
“Picks two smaller models — Mistral-seven-billion-Instruct, the dense seven-billion model, not to be confused with the bigger MIX-trul mixture-of-experts — and the two-billion Gemma instruction model.”Episode 004 — The Sycophancy Circuit That Survives Alignment Training