llm-architectures-explained
← /learn · 04

Dense and mixture of experts

Expert counts, active experts, shared experts and dense first layers.

Concept

In a dense model every token runs every weight. In a mixture of experts (MoE) the feed-forward block is replaced by many smaller feed-forward blocks, the experts, and a small router picks a few of them for each token. The model holds all the experts' weights but each token computes with only some of them, so the total parameter count (what the model knows, and what the memory must hold) and the active count (what each token costs) come apart. Switch Transformers made the point directly: "outrageous numbers of parameters, but a constant computational cost".

The design space has four knobs, and the interactive has a slider for each:

  • How many experts, and how many per token. Mixtral 8x7B routes each token to 2 of 8 experts, each as wide as a dense FFN. Among this data set's MoE models released since 2025, routing to 8 experts is by far the most common choice, out of pools of 8 to 896 experts.
  • Expert width (granularity). DeepSeekMoE split experts finely ("mNmN" narrow experts, "mKmK" of them active) so the router can combine them more flexibly. At the same active width, finer experts make many more combinations.
  • Shared experts, which every token uses, to hold common knowledge so the routed experts need not each relearn it (DeepSeekMoE again).
  • Dense first layers. DeepSeek-V3 keeps a dense FFN in its first three layers.

On the interactive's Llama 3 8B body (8.03B parameters dense), Mixtral's counts give 47.5B total and 13.7B active. DeepSeekMoE 16B's counts (64 experts an eighth of the width, 6 routed plus 2 shared) give 47.6B total but only 8.04B active: the same per-token work as the dense model, with six times the weights. DeepSeek-V3's counts (256 experts, 8 routed, 1 shared, 3 dense layers) at a sixteenth of the width give 85B total and 5.83B active, 6.9% of the total.

Loading the interactive…

Concept

What MoE does not save. At batch 1, a decode step reads only the active experts, so an MoE model decodes about as fast as a dense model of its active size. But all the experts must sit in memory (85B parameters in BF16 is 158 GiB), and a server batching many requests routes them to different experts, so each step reads far more than the active set. MoE trades memory and communication for compute; it shines where compute, not memory, is the limit. The compare view puts real models' totals and active counts side by side.

MoE: 88 models in the data set

Mixture of experts: each token uses a few of many feed-forward experts. Newest first; each links to its sourced page.

Shared expert: 55 models in the data set

One or more experts that every token uses. Newest first; each links to its sourced page.

Dense first layers: 45 models in the data set

The first layers keep a dense feed-forward block. Newest first; each links to its sourced page.

Latent MoE: 4 models in the data set

Experts work in a smaller latent width. Newest first; each links to its sourced page.

Maths

With EE routed experts of width ded_e, kk active, ss shared experts of width dsd_s, gated FFNs (three matrices) and model width dd, one MoE layer holds

Ptotal=3d (E de+s ds)+d EPactive=3d (k de+s ds)+d E,P_{\text{total}} = 3 d\,(E\,d_e + s\,d_s) + d\,E \qquad P_{\text{active}} = 3 d\,(k\,d_e + s\,d_s) + d\,E,

the last term being the router. A dense FFN of width dffd_{ff} holds 3d dff3 d\,d_{ff}; with de=dff/md_e = d_{ff}/m and k=mk = m routed experts the active FFN width equals the dense one, whatever EE is.

Code

// src/lib/chapters/model.ts — replace the dense FFN with experts (excerpt)
const dExpert = Math.floor(dense.d_ff! / granularity);
const moe: Spec["ffns"][string] = {
  type: "moe",
  experts,
  active,
  d_expert: dExpert,
  gated: dense.gated ?? true,
};

The presets' expert counts are the models' own (from their config files, on each model's page); the widths are this body's FFN divided by the granularity, not theirs.