Mixture of Experts (MoE)
A sparse architecture that routes each token to a subset of expert MLPs, increasing capacity without proportional compute cost.
Preparing spaced-repetition scheduler…
A Mixture of Experts (MoE) layer replaces a single dense MLP in a transformer block with N parallel expert MLPs and a small router that picks the top-k experts for each token.
For each input token x:
- The router computes logits R(x) ∈ ℝᴺ.
- The top-k entries (typically k=2) are kept; the rest are masked to -∞.
- A softmax over the kept entries produces a mixing weight.
- The output is the weighted sum of the chosen experts' outputs.
Why MoE:
- More capacity per unit of compute. A model with 8 experts, top-2 routing, activates only 2 experts per token. The expert layers have about 8× the parameters and 2× the active expert computation of one such MLP; shared layers, routing, and communication add overhead.
- Specialization. Different experts tend to specialize in different topics or syntactic patterns.
Trade-offs:
- Training instability. Routing decisions are discrete; load balancing and router stabilization can help; the method depends on the architecture.
- Memory cost. All experts must be resident or at least sharded across devices, which raises serving cost.
- Fine-tuning fragility. Routing can collapse to a few experts if not regularized.
MoE is widely used in modern open and proprietary frontier architectures as a way to scale capacity efficiently.
Key takeaways
- MoE routes each token to top-k experts out of N total.
- Parameter count grows but per-token compute grows much less.
- Routing needs load balancing; an auxiliary loss is one possible mechanism.
Self-check
What is the main efficiency claim of MoE?
What problem does a load-balancing loss address in MoE training?
References
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Original sparse gating research.
- Attention Is All You Need
Original transformer paper introducing scaled dot-product attention.
Last verified 2026-08-19.