Dense and mixture of experts
Expert counts, active experts, shared experts and dense first layers.
Concept
In a dense model every token runs every weight. In a mixture of experts (MoE) the feed-forward block is replaced by many smaller feed-forward blocks, the experts, and a small router picks a few of them for each token. The model holds all the experts' weights but each token computes with only some of them, so the total parameter count (what the model knows, and what the memory must hold) and the active count (what each token costs) come apart. Switch Transformers made the point directly: "outrageous numbers of parameters, but a constant computational cost".
The design space has four knobs, and the interactive has a slider for each:
- How many experts, and how many per token. Mixtral 8x7B routes each token to 2 of 8 experts, each as wide as a dense FFN. Among this data set's MoE models released since 2025, routing to 8 experts is by far the most common choice, out of pools of 8 to 896 experts.
- Expert width (granularity). DeepSeekMoE split experts finely ("" narrow experts, "" of them active) so the router can combine them more flexibly. At the same active width, finer experts make many more combinations.
- Shared experts, which every token uses, to hold common knowledge so the routed experts need not each relearn it (DeepSeekMoE again).
- Dense first layers. DeepSeek-V3 keeps a dense FFN in its first three layers.
On the interactive's Llama 3 8B body (8.03B parameters dense), Mixtral's counts give 47.5B total and 13.7B active. DeepSeekMoE 16B's counts (64 experts an eighth of the width, 6 routed plus 2 shared) give 47.6B total but only 8.04B active: the same per-token work as the dense model, with six times the weights. DeepSeek-V3's counts (256 experts, 8 routed, 1 shared, 3 dense layers) at a sixteenth of the width give 85B total and 5.83B active, 6.9% of the total.
Loading the interactive…
Concept
What MoE does not save. At batch 1, a decode step reads only the active experts, so an MoE model decodes about as fast as a dense model of its active size. But all the experts must sit in memory (85B parameters in BF16 is 158 GiB), and a server batching many requests routes them to different experts, so each step reads far more than the active set. MoE trades memory and communication for compute; it shines where compute, not memory, is the limit. The compare view puts real models' totals and active counts side by side.
MoE: 88 models in the data set
Mixture of experts: each token uses a few of many feed-forward experts. Newest first; each links to its sourced page.
- Kolibri-1
- Step 5 Preview
- Naive-N0.5-Flash
- MiMo-V2.6-Pro-RL
- MiMo-V2.6-Flash-RL
- Ember-1
- DeepSeek-V4.1-Flash
- Qwen3.8-Flash-Next
- Ornith 1.5 35B-A3B
- Nemotron 3.5 Lightning
- Ling 3.0 Flash
- Hy4 preview
- GLM-5.3-Flash
- Soofi S
- Solar Open 2
- Motif 3 Beta
- Laguna S 2.1
- Inkling
- North Mini Code 1.0
- Nemotron 3 Ultra
- MiniMax-M3
- Laguna XS 2.1
- Kimi K3
- Kimi K2.7 Code
- GLM-5.2
- ZAYA1-8B
- Mellum2 12B-A2.5B Thinking
- LFM2.5 8B-A1B
- Command A+
- Qwen3.6 35B-A3B
- MiniMax-M2.7
- MiMo-V2.5-Pro
- MiMo-V2.5
- Ling 2.6 1T
- Laguna XS.2
- Kimi K2.6
- Hy3 preview
- GLM-5.1
- DeepSeek-V4-Pro
- DeepSeek-V4-Flash
- Sarvam 30B
- Sarvam 105B
- Nemotron 3 Super
- Gemma 4 26B-A4B
- Step 3.5 Flash
- Qwen3.5 397B-A17B
- MiniMax-M2.5
- Ling 2.5 1T
- GLM-5
- Trinity Large
- Mistral Small 4
- LongCat-Flash-Lite
- Kimi K2.5
- Nemotron 3 Nano 30B-A3B
- MiMo-V2-Flash
- GLM-4.7
- DeepSeek-V3.2
- Mistral Large 3
- INTELLECT-3
- Gemini 3 Pro
- MiniMax-M2
- Kimi Linear 48B-A3B
- Qwen3-Next 80B-A3B
- Grok 2.5
- gpt-oss-20b
- gpt-oss-120b
- Qwen3-Coder 30B-A3B
- Kimi K2
- GLM-4.5-Air
- GLM-4.5
- Qwen3 30B-A3B
- Qwen3 235B-A22B
- Llama 4 Maverick
- Gemini 2.5 Pro
- MiniMax-Text-01
- DeepSeek-R1
- DeepSeek-V3
- OLMoE-1B-7B
- Mixtral 8x22B
- DeepSeek-V2
- Jamba v0.1
- Grok-1
- Qwen1.5-MoE-A2.7B
- Gemini 1.5 Pro
- DeepSeekMoE 16B
- Mixtral 8x7B
- GLaM (64B/64E)
- Switch-C
Shared expert: 55 models in the data set
One or more experts that every token uses. Newest first; each links to its sourced page.
- Kolibri-1
- Ember-1
- DeepSeek-V4.1-Flash
- Qwen3.8-Flash-Next
- Ornith 1.5 35B-A3B
- Nemotron 3.5 Lightning
- Ling 3.0 Flash
- Hy4 preview
- GLM-5.3-Flash
- Soofi S
- Solar Open 2
- Motif 3 Beta
- Laguna S 2.1
- Inkling
- Nemotron 3 Ultra
- MiniMax-M3
- Laguna XS 2.1
- Kimi K3
- Kimi K2.7 Code
- GLM-5.2
- Command A+
- Qwen3.6 35B-A3B
- Ling 2.6 1T
- Laguna XS.2
- Kimi K2.6
- Hy3 preview
- GLM-5.1
- DeepSeek-V4-Pro
- DeepSeek-V4-Flash
- Sarvam 30B
- Sarvam 105B
- Nemotron 3 Super
- Step 3.5 Flash
- Qwen3.5 397B-A17B
- Ling 2.5 1T
- GLM-5
- Trinity Large
- Mistral Small 4
- Kimi K2.5
- Nemotron 3 Nano 30B-A3B
- GLM-4.7
- DeepSeek-V3.2
- Mistral Large 3
- INTELLECT-3
- Kimi Linear 48B-A3B
- Qwen3-Next 80B-A3B
- Kimi K2
- GLM-4.5-Air
- GLM-4.5
- Llama 4 Maverick
- DeepSeek-R1
- DeepSeek-V3
- DeepSeek-V2
- Qwen1.5-MoE-A2.7B
- DeepSeekMoE 16B
Dense first layers: 45 models in the data set
The first layers keep a dense feed-forward block. Newest first; each links to its sourced page.
- Naive-N0.5-Flash
- MiMo-V2.6-Pro-RL
- MiMo-V2.6-Flash-RL
- Ember-1
- Ling 3.0 Flash
- Hy4 preview
- GLM-5.3-Flash
- Motif 3 Beta
- Laguna S 2.1
- Inkling
- North Mini Code 1.0
- MiniMax-M3
- Laguna XS 2.1
- Kimi K3
- Kimi K2.7 Code
- GLM-5.2
- LFM2.5 8B-A1B
- MiMo-V2.5-Pro
- MiMo-V2.5
- Ling 2.6 1T
- Laguna XS.2
- Kimi K2.6
- Hy3 preview
- GLM-5.1
- Sarvam 30B
- Sarvam 105B
- Step 3.5 Flash
- Ling 2.5 1T
- GLM-5
- Trinity Large
- LongCat-Flash-Lite
- Kimi K2.5
- MiMo-V2-Flash
- GLM-4.7
- DeepSeek-V3.2
- Mistral Large 3
- INTELLECT-3
- Kimi Linear 48B-A3B
- Kimi K2
- GLM-4.5-Air
- GLM-4.5
- DeepSeek-R1
- DeepSeek-V3
- DeepSeek-V2
- DeepSeekMoE 16B
Latent MoE: 4 models in the data set
Experts work in a smaller latent width. Newest first; each links to its sourced page.
Maths
With routed experts of width , active, shared experts of width , gated FFNs (three matrices) and model width , one MoE layer holds
the last term being the router. A dense FFN of width holds ; with and routed experts the active FFN width equals the dense one, whatever is.
Code
// src/lib/chapters/model.ts — replace the dense FFN with experts (excerpt)
const dExpert = Math.floor(dense.d_ff! / granularity);
const moe: Spec["ffns"][string] = {
type: "moe",
experts,
active,
d_expert: dExpert,
gated: dense.gated ?? true,
};
The presets' expert counts are the models' own (from their config files, on each model's page); the widths are this body's FFN divided by the granularity, not theirs.