Glossary · Term

mixture-of-experts

← all terms

Definition

Plain language

A model design where only a fraction of the parameters fire on any one input, letting models be very large but cheap to run.

As stated in the literature

A neural network architecture in which only a sparse subset of expert sub-networks is activated per token, enabling large total parameter counts at lower per-token compute.

Also called: MoE, mixture of experts, sparse mixture-of-experts

Why it matters: It lets models grow in total knowledge without proportionally growing the compute needed per token at inference.

For example, a 200-billion-parameter MoE model might only activate ~10 billion parameters when answering a single question.

Heard on the show

“Picks two smaller models — Mistral-seven-billion-Instruct, the dense seven-billion model, not to be confused with the bigger MIX-trul mixture-of-experts — and the two-billion Gemma instruction model.”
Episode 004 — The Sycophancy Circuit That Survives Alignment Training

Related terms