T2.6 — What Mixture of Experts Changes¶
In one line: MoE activates only a few expert MLPs per token, reducing arithmetic relative to total parameters while retaining the memory and communication cost of a much larger model.
| Skill | 00-foundations |
| Module | M2 — Transformer Internals from the Inference Angle |
| Audience | New graduate engineer |
| Time | ~110 min |
| Prereqs | T2.1; softmax and top-k selection |
| Status | built |
Why this matters¶
A model advertised as “30B total, 3B active” does not behave like a dense 3B model. The active count describes token-level arithmetic; the total count still influences weight capacity, loading time, and distributed placement, while uneven routing can leave some GPUs overloaded and others idle.
The mental model¶
A dense decoder sends every token through the same MLP. A mixture-of-experts layer owns several alternative MLPs called experts. A small router scores the token, selects the top k experts, and combines their outputs.

The router selects a few expert paths for each token and combines their outputs.
Think of a hospital with many specialist departments. The building must contain every department, but one patient visits only the few specialists selected for that case. Total departments determine building capacity; visited departments determine work per patient. MoE separates total parameters from active parameters in the same way.
flowchart LR
T[Token] --> R{{Router scores}}
R -->|top-k| E1[[Expert 1]]
R -->|top-k| E4[[Expert 4]]
R -.not selected.-> E2[[Expert 2]]
E1 --> C[Weighted combine]
E4 --> C
Sparse activation reduces expert arithmetic per token, not the number of expert weights the system must store.
The mechanism¶
For token vector x, the router produces logits over E experts and converts them into routing weights. Top-k selection keeps the k chosen experts:
Each E_e is usually a gated MLP similar to T2.1. Some architectures also include shared experts that process every token. Router parameters are small compared with expert weights, but router decisions determine communication and load balance.
Top-k routing follows four conceptual steps. The router produces one score per expert. A normalization produces routing weights. Top-k selects expert indices. Finally, selected expert outputs are weighted and summed. Implementations group tokens by expert so each expert receives a useful batch rather than launching a tiny operation for every token.
Total and active parameters answer different questions¶
Suppose eight experts each contain 100 million parameters and top-2 routing selects two per token. Expert storage is 800 million parameters, while expert arithmetic per token touches about 200 million. Add shared attention, embeddings, router, and any shared experts to both totals as appropriate.
The toy artifact makes this distinction explicit. With eight small experts and top-2 routing, about 25.1% of its parameters are active per token. Yet all 1.05 million parameters exist in memory. A production “30B total, 3B active” model similarly requires weight capacity closer to its total size even though token-level expert compute resembles the smaller active count.

All expert weights occupy memory even though each token activates only a few.
Weight bandwidth complicates the picture. If a batch routes tokens across every expert, the system may read many or all expert weights during the step even though each token uses only two. Expert locality and batch composition therefore matter. Sparse FLOPs do not guarantee proportional latency reduction.
“Active parameters” is usually a per-token statement. Different tokens in one batch may choose different experts, so the union of weights accessed by the batch can include every expert. A top-2 model can therefore touch all expert weights during a busy step even though no individual token does.
Routing creates an assignment problem¶
The router may send more tokens to some experts. The artifact counts every token–expert assignment and reports maximum load divided by mean load. Its seeded run produced a ratio of about 1.09: the busiest expert received roughly 9% more assignments than the average. A trained router or skewed workload can be far less balanced.

Uneven routing makes the busiest expert the step’s critical path.
Capacity limits often cap how many tokens one expert accepts. Overflow may be dropped, rerouted, or handled by a fallback depending on architecture and training. Auxiliary load-balancing losses encourage even use during training, but serving traffic can still shift distributions.
Router confidence matters alongside counts. Weights 0.51 and 0.49 make both selected experts important; 0.99 and 0.01 spend almost the same compute for a very different contribution. Architectures normalize and bias these weights differently, so inspect the implementation before interpreting router logits.
Distributed MoE adds all-to-all communication¶
With expert parallelism, different GPUs own different experts. Tokens begin on the GPU holding their current hidden states, then must travel to the GPUs holding selected experts. After expert computation, outputs travel back. This dispatch and return commonly uses all-to-all communication.
Communication cost depends on tokens, hidden size, dtype, expert placement, and routing balance. One overloaded expert can become the step’s critical path while other GPUs wait. Unlike tensor parallelism’s regular collective pattern, MoE traffic depends on routing decisions made from the data.
Suppose four GPUs each own two experts. A token on GPU 0 may select experts on GPUs 1 and 3, sending its hidden vector to both. Both outputs then return for combination. Repeated for every token and MoE layer, this becomes a major fabric workload.
In practice¶
Read total parameters, active experts, top-k, expert intermediate size, shared-expert count, and placement. Do not compare an MoE’s active parameter label directly with a dense model’s total label without explaining memory and communication.
Monitor per-expert token counts, maximum-to-mean load, capacity overflow, all-to-all time, and expert compute time. Aggregate GPU utilization can hide one hot rank and several waiting ranks.
Batch composition affects locality. Routing many tokens to a small expert subset can reuse weights but overload those experts; broad routing balances assignments but may fetch more expert weights. Measure with the actual domain because router distributions are input-dependent.
Compare MoE and dense models with separate columns for total weight bytes, active FLOPs, achieved batch, communication, latency, throughput, and quality. One parameter label cannot represent all of these costs.
Failure modes¶
| Symptom | Cause | Fix |
|---|---|---|
| “3B active” model does not fit like a dense 3B | Total expert weights still occupy memory | Size capacity from total stored parameters plus runtime state |
| GPUs show uneven step times | Router load or expert placement is imbalanced | Inspect per-expert counts and placement; tune capacity/balancing strategy |
| Sparse FLOP count yields little latency benefit | Weight reads, small expert GEMMs, routing, or all-to-all dominates | Profile each stage and increase useful token grouping where SLO permits |
| Tokens are dropped or quality changes under load | Expert capacity is exceeded | Inspect overflow policy and capacity factor; reduce batch or rebalance |
| Toy random-router balance is treated as production evidence | Real router and domain distribution differ | Measure trained weights on representative requests |
Do it¶
Run the toy MoE artifact. Change expert count, top-k, token count, and random seed. Predict active fraction, then observe how load varies even when average assignments are fixed.
Success means explaining total parameters, active parameters, and assignment imbalance as three separate facts. Extend the experiment with expert placement before drawing distributed performance conclusions.
Check¶
- Why does top-2 routing reduce token-level expert arithmetic without reducing stored expert weights?
- What does maximum-to-mean expert load reveal that average load does not?
- Why can expert parallelism require data-dependent all-to-all communication?