Long-context techniques
RoPE scaling, windows, compressed and sparse attention, and hybrids, at 1M tokens.
Concept
Reading a million tokens poses two separate problems.
Can the model use positions it never trained on? That is a positional encoding question. Raising RoPE's base, interpolating positions, YaRN's per-wavelength rescaling (chapter 2) and NoPE layers all answer it, and none of them changes what a token costs.
Can anyone afford to run it? That is an attention question. With full attention in every layer, the cache grows by the same amount for every token, and every new token reads all of it. Llama 3.1 405B, GQA in all 126 layers, holds 63 GiB of cache per 128K-token sequence in BF16; at 1M tokens it would need 504 GiB for one sequence, and 9.47 TFLOP per generated token, 91.5% of it attention. The answers are the mixers of chapter 1, now on real models:
- Windows. Gemma 3 27B's five sliding-window layers per global layer keep its 1M-token cache to 80.4 GiB.
- Compress the cache. DeepSeek-V3's MLA latent holds 68.6 GiB at 1M, but every token still reads all of it: 5.31 TFLOP per token, 98.6% attention in this cost model's expanded form.
- Read less of it. DeepSeek-V3.2 adds an indexer that picks 2,048 cached tokens per query; at 1M tokens its decode drops to 1.13 TFLOP.
- Replace most layers with a fixed state. Qwen3-Next (Gated DeltaNet, 3:1) needs 24 GiB at 1M; MiniMax-Text-01 (Lightning attention, 7:1) 40.1 GiB.
- Compress and sparsify together. DeepSeek V4-Flash compresses keys and values along the sequence and lets an indexer choose among the compressed entries: 6.72 GiB at 1M tokens, the smallest here.
Loading the interactive…
Concept
Things to try. Switch the chart to decode FLOPs: every line starts flat, where the weights dominate (two FLOPs per active parameter), and bends upward once attention takes over. The windowed, hybrid and sparse models bend later and stay lower. Then look at the last column: at 1M tokens even the most frugal of these spends most of each decode step on token mixing, not on the feed-forward weights that the parameter count measures.
Past a model's supported context length these numbers are what its architecture would cost, not a claim that it works there: each model's page gives the context its makers state.
Sliding window: 35 models in the data set
Some layers attend only to the most recent W tokens. Newest first; each links to its sourced page.
- Kolibri-1
- Naive-N0.5-Flash
- MiMo-V2.6-Pro-RL
- MiMo-V2.6-Flash-RL
- DeepSeek-V4.1-Flash
- Muse Glimmer 30B
- Motif 3 Beta
- Laguna S 2.1
- Inkling
- North Mini Code 1.0
- Laguna XS 2.1
- Mellum2 12B-A2.5B Thinking
- Gemma 4 12B
- Command A+
- MiMo-V2.5-Pro
- MiMo-V2.5
- Laguna XS.2
- DeepSeek-V4-Pro
- DeepSeek-V4-Flash
- Gemma 4 E4B
- Gemma 4 E2B
- Gemma 4 31B
- Gemma 4 26B-A4B
- Tiny Aya
- Step 3.5 Flash
- Trinity Large
- MiMo-V2-Flash
- Olmo 3 32B
- Olmo 3 7B
- gpt-oss-20b
- gpt-oss-120b
- Gemma 3 270M
- Gemma 3 27B
- Gemma 2 27B
- Mistral 7B
Compressed KV: 3 models in the data set
Keys and values compressed along the sequence (DeepSeek V4 CSA/HCA). Newest first; each links to its sourced page.
Linear attention: 20 models in the data set
A fixed-size recurrent state instead of a growing cache. Newest first; each links to its sourced page.
Mamba (SSM): 9 models in the data set
Selective state-space layers. Newest first; each links to its sourced page.
Maths
For a model whose layers are full attention (), windowed (, window ) or recurrent (, state ), the cache for a context of tokens is
with values per token in layer and , bytes per value. Only the first term grows without bound, so for long the cache is set by the number of full-attention layers and their , which is what MLA ( against GQA's ) and DeepSeek V4's compression shrink. Decode FLOPs follow the entries each query reads: in , in , the indexer's top- with sparse attention.
Code
// src/lib/arch/costModel.ts — growing, windowed and fixed caches (excerpt)
if (e > 0) {
if (w > 0) {
windowed = windowed + e * Math.min(context, w);
} else {
growing = growing + e * context;
perTokenUnbounded = perTokenUnbounded + e;
}
}
state = state + stateElems(spec, run);