The attention family
MHA, GQA, MQA, MLA, sliding windows, sparse attention, and linear and state-space hybrids.
Concept
The token mixer is the part of a block that lets a position read earlier positions. Since 2017 it has been the place where architectures differ most, and nearly every change has the same target: the cache a model keeps per sequence while it generates, and the work each new token does against it.
- Multi-head attention (MHA), the original Transformer's: every query head has its own key and value head, so each layer caches two vectors per head per token.
- Multi-query attention (MQA) keeps all the query heads but shares one key and one value head between them (Shazeer, 2019). PaLM used it.
- Grouped-query attention (GQA) sits between the two: query heads in groups, one key/value head per group (Ainslie et al., 2023). Llama 3 8B has 32 query heads and 8 key/value heads.
- Multi-head latent attention (MLA), from DeepSeek-V2, caches one compressed latent vector per token per layer, plus a small separate key for positions, and rebuilds every head's keys and values from it.
- Sliding windows make most layers attend only to the most recent W tokens. Gemma 3 interleaves five local layers (a 1,024-token window) with one global layer.
- Sparse attention keeps the whole cache but lets a cheap indexer pick which cached tokens each query reads (DeepSeek-V3.2 picks 2,048).
- Linear attention and state-space layers (Gated DeltaNet, Lightning attention, Mamba) replace the growing cache with a fixed-size state. They are almost always mixed with some full-attention layers: Qwen3-Next uses three Gated DeltaNet layers per attention layer, Jamba seven Mamba layers per attention layer.
The interactive puts all eight on one body: Llama 3 8B's width, depth, feed-forward blocks and vocabulary, with only the token mixer swapped. So every difference you see is the mixer's alone. At a 128K-token context in BF16, GQA's cache is 16 GiB per sequence; MHA's is 4× that, 64 GiB; MQA's 2 GiB; MLA's 4.5 GiB. The sliding-window mix holds 2.61 GiB, and the Mamba-2 hybrid 2.06 GiB, almost all of it in its four global layers.
Loading the interactive…
Concept
Things to try.
- Look at the chart from 1K to 1M. Every line ends up rising in step with the context, because every variant keeps some layers whose cache grows. The window, linear and state-space hybrids start almost flat, where their fixed-size part is all there is, and then rise from a much lower level, because only one layer in four, six or eight keeps a growing cache.
- Compare decode FLOPs at 128K. The attention term is proportional to the tokens each query reads, so the sliding-window mix needs 26.2 GFLOP per token against GQA's 83.7. MLA needs more (101 GFLOP) in this cost model, which counts its attention in the expanded form: MLA saves memory, not arithmetic.
- Switch the cache to FP8. Every cache halves; the recurrent state is counted at its own precision.
GQA: 86 models in the data set
Grouped-query attention: several query heads share one key/value head. Newest first; each links to its sourced page.
- Kolibri-1
- Naive-N0.5-Flash
- MiMo-V2.6-Pro-RL
- MiMo-V2.6-Flash-RL
- Qwen3.8-Flash-Next
- Qwen3.8 27B
- Ornith 1.5 35B-A3B
- Nemotron 3.5 Lightning
- Muse Glimmer 30B
- Soofi S
- Solar Open 2
- Nanbeige4.2 3B
- Laguna S 2.1
- Inkling
- BTL-3
- Antares 1B
- VibeThinker-3B
- North Mini Code 1.0
- Nemotron 3 Ultra
- MiniMax-M3
- Laguna XS 2.1
- ZAYA1-8B
- Mellum2 12B-A2.5B Thinking
- LFM2.5 8B-A1B
- Gemma 4 12B
- Command A+
- Qwen3.6 35B-A3B
- Qwen3.6 27B
- MiniMax-M2.7
- MiMo-V2.5-Pro
- MiMo-V2.5
- Laguna XS.2
- Hy3 preview
- Granite 4.1 30B
- Sarvam 30B
- Nemotron 3 Super
- Nemotron 3 Nano 4B
- LFM2.5 350M
- Gemma 4 E4B
- Gemma 4 31B
- Gemma 4 26B-A4B
- Tiny Aya
- Step 3.5 Flash
- Qwen3.5 397B-A17B
- Nanbeige4.1 3B
- MiniMax-M2.5
- Trinity Large
- LFM2.5 1.2B
- Nemotron 3 Nano 30B-A3B
- MiMo-V2-Flash
- GLM-4.7
- Olmo 3 32B
- INTELLECT-3
- MiniMax-M2
- Qwen3-Next 80B-A3B
- Grok 2.5
- gpt-oss-20b
- gpt-oss-120b
- Qwen3-Coder 30B-A3B
- GLM-4.5-Air
- GLM-4.5
- SmolLM3 3B
- Qwen3 8B
- Qwen3 4B
- Qwen3 32B
- Qwen3 30B-A3B
- Qwen3 235B-A22B
- Qwen3 0.6B
- Llama 4 Maverick
- Mistral Small 3.1
- Gemma 3 27B
- MiniMax-Text-01
- Phi-4
- Qwen2.5 72B
- Llama 3.2 3B
- Llama 3.2 1B
- Llama 3.1 405B
- Gemma 2 27B
- Mixtral 8x22B
- Llama 3 8B
- Jamba v0.1
- Grok-1
- Mixtral 8x7B
- Mistral 7B
- Llama 2 70B
- Falcon-40B
MLA: 24 models in the data set
Multi-head latent attention: keys and values cached as one low-rank latent vector. Newest first; each links to its sourced page.
Sliding window: 35 models in the data set
Some layers attend only to the most recent W tokens. Newest first; each links to its sourced page.
- Kolibri-1
- Naive-N0.5-Flash
- MiMo-V2.6-Pro-RL
- MiMo-V2.6-Flash-RL
- DeepSeek-V4.1-Flash
- Muse Glimmer 30B
- Motif 3 Beta
- Laguna S 2.1
- Inkling
- North Mini Code 1.0
- Laguna XS 2.1
- Mellum2 12B-A2.5B Thinking
- Gemma 4 12B
- Command A+
- MiMo-V2.5-Pro
- MiMo-V2.5
- Laguna XS.2
- DeepSeek-V4-Pro
- DeepSeek-V4-Flash
- Gemma 4 E4B
- Gemma 4 E2B
- Gemma 4 31B
- Gemma 4 26B-A4B
- Tiny Aya
- Step 3.5 Flash
- Trinity Large
- MiMo-V2-Flash
- Olmo 3 32B
- Olmo 3 7B
- gpt-oss-20b
- gpt-oss-120b
- Gemma 3 270M
- Gemma 3 27B
- Gemma 2 27B
- Mistral 7B
Sparse attention: 10 models in the data set
An indexer picks which cached tokens each query reads. Newest first; each links to its sourced page.
Hybrid: 27 models in the data set
More than one kind of token mixer in the stack. Newest first; each links to its sourced page.
- Ember-1
- Qwen3.8-Flash-Next
- Qwen3.8 27B
- Ornith 1.5 35B-A3B
- Nemotron 3.5 Lightning
- Ling 3.0 Flash
- GLM-5.3-Flash
- Soofi S
- Solar Open 2
- BTL-3
- Nemotron 3 Ultra
- Kimi K3
- LFM2.5 8B-A1B
- Qwen3.6 35B-A3B
- Qwen3.6 27B
- Ling 2.6 1T
- Nemotron 3 Super
- Nemotron 3 Nano 4B
- LFM2.5 350M
- Qwen3.5 397B-A17B
- Ling 2.5 1T
- LFM2.5 1.2B
- Nemotron 3 Nano 30B-A3B
- Kimi Linear 48B-A3B
- Qwen3-Next 80B-A3B
- MiniMax-Text-01
- Jamba v0.1
Maths
Per token, summed over layers, a softmax-attention cache holds
| Mixer | Values per token |
|---|---|
| MHA | |
| GQA | |
| MQA | |
| MLA | |
| Window | , for the last tokens only |
for query heads of dimension , key/value heads, a compressed latent of values and a decoupled rotary key of . For the Llama 3 8B body (, , , ) in BF16 that is 512 KiB per token for MHA, 128 KiB for GQA, 16 KiB for MQA and, with DeepSeek-V3's latent sizes (, ), 36 KiB for MLA.
A linear-attention or state-space layer instead keeps a state of fixed size: for a Gated DeltaNet layer with value heads, for a Mamba-2 layer with state size , plus a few tokens of convolution history. That is why the hybrids' curves flatten.
Decode FLOPs per token at context are for the matrix multiplications, plus, per attention layer, for each cached token the query reads: for full attention, in a window, and the indexer's top- for sparse attention (whose indexer still scores all tokens with FLOPs each).
Code
// src/lib/arch/costModel.ts — the cache one layer adds per token (excerpt)
if (t === "attn") {
const kv = m.kv_heads!;
const hd = m.head_dim!;
const vd = m.v_head_dim || m.head_dim!;
let e = m.k_eq_v ? kv * hd : kv * (hd + vd);
The variants are built in reference/chapter_model.py (attention_variants)
and costed by the cost model in Python and in TypeScript;
tests/unit/chapters/model.test.ts asserts that every number in the table
is identical in the two.