Architectureadvanced6 min read

Mixture of Experts (MoE)

A sparse architecture that routes each token to a subset of expert MLPs, increasing capacity without proportional compute cost.

Preparing spaced-repetition scheduler…

A Mixture of Experts (MoE) layer replaces a single dense MLP in a transformer block with N parallel expert MLPs and a small router that picks the top-k experts for each token.

For each input token x:

  1. The router computes logits R(x) ∈ ℝᴺ.
  2. The top-k entries (typically k=2) are kept; the rest are masked to -∞.
  3. A softmax over the kept entries produces a mixing weight.
  4. The output is the weighted sum of the chosen experts' outputs.

Why MoE:

  • More capacity per unit of compute. A model with 8 experts, top-2 routing, activates only 2 experts per token. The expert layers have about 8× the parameters and 2× the active expert computation of one such MLP; shared layers, routing, and communication add overhead.
  • Specialization. Different experts tend to specialize in different topics or syntactic patterns.

Trade-offs:

  • Training instability. Routing decisions are discrete; load balancing and router stabilization can help; the method depends on the architecture.
  • Memory cost. All experts must be resident or at least sharded across devices, which raises serving cost.
  • Fine-tuning fragility. Routing can collapse to a few experts if not regularized.

MoE is widely used in modern open and proprietary frontier architectures as a way to scale capacity efficiently.

MoE routing for a single token. The router picks the top experts; the rest are skipped. The output is a weighted combination of the activated experts.

Key takeaways

  • MoE routes each token to top-k experts out of N total.
  • Parameter count grows but per-token compute grows much less.
  • Routing needs load balancing; an auxiliary loss is one possible mechanism.

Self-check

  1. What is the main efficiency claim of MoE?

  2. What problem does a load-balancing loss address in MoE training?

References

  1. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer — Shazeer et al. / arXiv

    Original sparse gating research.

  2. Attention Is All You Need — NeurIPS / arXiv

    Original transformer paper introducing scaled dot-product attention.

Last verified 2026-08-19.