← All posts

Only 3% of the Brain Wakes Up: The Mixture-of-Experts Playbook Behind September's Biggest Models

2026-09-26 · ML architectures (mixture-of-experts)

Two open-weight releases, twelve days apart. DeepSeek V4.1 Flash (September 10): 552 billion parameters — and about 8 billion of them fire on each input token, 16 billion on output. Xiaomi MiMo-V2.6 Pro (September 22): 1.02 trillion parameters, with just 42 billion active per token. Both MIT-licensed. Both cheaper per token than models a tenth their size.

The trick isn't new. It's a 2017 idea (Shazeer et al.'s sparsely-gated MoE, 137B parameters) that quietly conquered the frontier: of the seven frontier open-weight models people actually deploy in 2026, six are Mixture-of-Experts. This is the architecture eating the scaling curve — so here's how it works, where it breaks, and what September's research says about fixing it.

ELI5: the hospital with specialists on call

A dense transformer is a hospital where every doctor examines every patient. Chest pain? All 200 doctors walk in. Splinter? All 200 doctors walk in. You pay for all of them, every time.

A Mixture-of-Experts model is a hospital with a triage nurse. Each patient (token) gets routed to only the 2–8 specialists relevant to their case. The hospital still employs 200 specialists — that's the model's knowledge capacity — but each patient only pays for a handful. Total parameters are the staff directory. Active parameters are the bill.

That's the whole bargain: a huge brain, a small compute bill. DeepSeek-V3 carries 671B parameters but computes like a 37B model — an 18× gap between what it knows and what it spends. The September releases pushed it further: V4.1 Flash runs 552B of knowledge on ~16B of compute per output token, a 34× gap. Only ~3% of the brain wakes up.

Diagram 1

How it works: a router, k experts, and a weighted vote

In a standard transformer block, the feed-forward network (FFN) holds most of the parameters and most of the compute. MoE replaces that single FFN with N smaller expert networks plus a router (the gating network):

  1. Score. The router reads the token and scores every expert: scores = softmax(x · W_router).
  2. Select. Keep only the top-k (typically 2–8). The rest sit idle for this token.
  3. Combine. Each selected expert processes the token; outputs merge as a weighted sum using the router's scores: y = Σ gᵢ · Expertᵢ(x).

Modern variants add a shared expert that every token also passes through (DeepSeek's design: 256 routed experts + 1 shared), which soaks up common knowledge so routed experts can specialize.

Training has one notorious failure mode: router collapse. Left alone, the router discovers 3 favorite experts and starves the rest — 61 experts never learn anything. The standard fix is an auxiliary load-balancing loss (from Switch Transformer): penalize the router when traffic concentrates, L_aux = α · N · Σ fᵢ · Pᵢ, where fᵢ is the fraction of tokens sent to expert i and Pᵢ the average router probability. It's a tax on favoritism.

And the catch the marketing skips: the savings are compute, not memory. Every expert must sit in VRAM whether or not a given token uses it. DeepSeek-V3's 671B parameters in FP8 need ~671 GB of weights — a multi-GPU server — even though each token computes like a 37B model. MoE buys you a bigger brain for the same FLOP bill, then hands you the memory bill.

Diagram 2

The numbers (SOTA section)

September's leaderboard. The frontier is now an MoE monoculture, and the licenses are getting friendlier:

Model Released Total / active Experts (top-k) License
MiMo-V2.6 Pro Sep 22, 2026 1.02T / 42B 384 (8) MIT
DeepSeek V4.1 Flash Sep 10, 2026 552B / 8B in, 16B out — MIT
GLM-5.3-Flash Aug 26, 2026 320B / 18B — MIT
Qwen3-235B-A22B 2025 235B / 22B 128 (8) open
DeepSeek-V3 Dec 2024 671B / 37B 256+1 shared (8) open
Instella-MoE-16B-A3B Jul 2026 16B / 2.8B 64 routed (6) + 2 shared fully open*

*AMD's Instella is the "fully open" outlier: trained end-to-end on MI300X/MI325X GPUs with open Primus/Miles frameworks — 12.7% faster training and 39.2% lower time-to-first-token from its Gated MLA + FarSkip-Collective design. (Weights carry a ResearchRAIL non-commercial limit; the training code is MIT.)

Pricing tells the story: V4.1 Flash at $0.15/$0.60 per million tokens off-peak and MiMo-V2.6 Pro at $0.435/$0.87 — frontier-adjacent intelligence at dense-model prices, because the provider's compute cost is set by the active count.

The research front is attacking MoE's three chronic diseases:

Diagram 3

The serving tax. Because no single GPU holds all experts, MoE serving shards them across GPUs — expert parallelism — and each token travels to the GPUs holding its experts, then the weighted outputs travel back. That all-to-all shuffle is pure overhead: bandwidth and stragglers cap throughput. Worse, it creates a verification blind spot: a June 2026 analysis ("The Expert Shuffle") showed a decentralized provider can silently cut top-k from 8 to 4 ("k-shunting") and stay nearly undetectable — because the Standing Committee dominates both runs, outputs overlap. Sampling-based checks can't catch it; only measurement of intermediate computation can.

Diagram 4

The limits, stated plainly

  1. Memory doesn't sparsify. You pay VRAM for every parameter and FLOPs for the active few. "Open weights" at 552B–1T params means a multi-GPU server (V4.1 Flash's checkpoint is ~510 GB).
  2. Routing is the fragile part. Collapse, imbalance, and the standing-committee effect mean much of the expert count is decorative. UniPool's random-routing result is humbling: deep routers barely matter.
  3. Serving is a distributed-systems problem. All-to-all dispatch/combine, capacity scheduling, stragglers — MoE inference engineering looks more like HPC than like running a dense model.
  4. Verification has a hole. k-shunting in decentralized serving is currently hard to detect from outputs alone. If you're buying MoE inference from a third party, ask what they measure.

Takeaways

  1. MoE decouples knowledge from compute. Total params = capacity (memory bill); active params = cost per token (compute bill). September's releases push the gap to 24–34×.
  2. The router is the whole game. Top-k softmax gating + weighted sum is simple; keeping all experts alive via load-balancing losses is the hard part — and 2026 research (LPR, UniPool) is winning.
  3. Deep experts are redundant. UniPool showed random routing in deep layers costs ~1 point — the future is shared expert pools, not per-layer expert farms.
  4. The frontier is an MoE monoculture with MIT licenses. 6 of 7 deployed frontier open-weight models are MoE, and the September flagships ship MIT weights. The moat moved from weights to serving.
  5. Watch the serving layer, not the parameter count. All-to-all communication and k-shunting mean the interesting MoE problems in 2026 are systems and verification, not scale.

Companion notebook: moe_mixture_of_experts.ipynb — build a tiny MoE layer from scratch in NumPy (router, top-2 gating, load-balancing aux loss), train it on a toy task, and watch router collapse happen without the aux loss and disappear with it.

Sources: DeepSeek V4.1 Flash release details · MiMo-V2.6 Pro release · GLM-5.3-Flash model card · The Expert Shuffle: MoE serving's verification problem · The Expert Gap: routing and the verification blind spot · What is Mixture of Experts? (Sep 2026) · UniPool paper wiki · Latent Prototype Routing (LPR) · GPT-OSS-20B deployment analysis · AMD Instella-MoE

Companion notebook

moe_mixture_of_experts.ipynb — the runnable tutorial for this post (download, or open it in Colab/Jupyter).

← All posts