Concept
Attention on its own is blind to order: shuffle the cached keys and values and each query's output is unchanged. Position has to be put in somewhere, and the choices have changed more than once.
- Sinusoidal (the original Transformer) and learned absolute embeddings (GPT-2, GPT-3, BERT, OPT) add a position vector to each token's embedding. A learned table has one row per position, so it also fixes the longest sequence the model can read.
- Relative biases (T5) and ALiBi (BLOOM) add nothing to the embeddings; they add a bias to every attention score that depends on the distance between query and key. ALiBi's is a fixed linear penalty per head, chosen so a model trained on short sequences can read longer ones (Press et al., 2021).
- Rotary position embedding (RoPE) (Su et al., 2021) rotates each query and key, pair of dimensions by pair of dimensions, through an angle proportional to its position. The dot product of a rotated query and key then depends only on how far apart they are. Almost every open model since LLaMA uses it.
- Partial RoPE rotates only some of each head's dimensions and leaves the rest position-free: GLM-4.5 rotates half, Qwen3-Next a quarter. MLA (chapter 1) carries position in a small separate rotary key, because a rotation cannot pass through its compressed latent.
- NoPE layers use no positional encoding at all; the causal mask alone lets a decoder infer order (Kazemnejad et al., 2023). Some recent models interleave NoPE layers with RoPE layers.
RoPE has one knob that matters for long context: its base θ. Pair i of a head with d rotated dimensions turns through an angle of radians per token, so its wavelength is tokens. The first pair spins every 6.3 tokens; the last ones turn so slowly that they barely move across the whole context. A pair that turns less than once across the context still carries information, but only in its slow drift, and those are the dimensions that misbehave when a model reads past its training length. Raising θ stretches every wavelength. Mistral 7B's θ = 10,000 leaves 4 of its 64 pairs slower than a 32K context; Qwen3 8B's θ = 1,000,000 leaves 24.
Loading the interactive…
Concept
Things to try.
- Pick Mistral 7B, then Qwen3 8B, at 32K. Same 128-dimension heads, wavelengths a hundred times longer at the slow end.
- Pick GLM-4.5 or Qwen3-Next. Partial RoPE halves (or quarters) the pairs, and the remaining dimensions carry no position at all.
- Pick DeepSeek-V3. Its rotary key is only 64 dimensions wide, with θ = 10,000: at 128K no pair is slower than the context.
Extending context after training. Position Interpolation (Chen et al., 2023) squeezes longer inputs into the trained range by scaling positions down; YaRN (Peng et al., 2023) sorts the pairs by exactly the comparison the interactive draws: a pair whose wavelength is much shorter than the trained context is left alone, one at least as long as the context is interpolated, and those in between get a blend (section 3.2, p. 5). Neither changes the cache, the parameters or the FLOPs: positional encoding costs almost nothing to run. What it changes is how far a trained model can read, which chapter 6 takes up.
RoPE: 126 models in the data set
Rotary position embeddings. Newest first; each links to its sourced page.
- Kolibri-1
- Naive-N0.5-Flash
- MiMo-V2.6-Pro-RL
- MiMo-V2.6-Flash-RL
- Ember-1
- DeepSeek-V4.1-Flash
- Qwen3.8-Flash-Next
- Qwen3.8 27B
- Ornith 1.5 35B-A3B
- Nemotron 3.5 Lightning
- Muse Glimmer 30B
- Ling 3.0 Flash
- Hy4 preview
- GLM-5.3-Flash
- Soofi S
- Solar Open 2
- Nanbeige4.2 3B
- Motif 3 Beta
- Laguna S 2.1
- Inkling
- BTL-3
- Antares 1B
- VibeThinker-3B
- North Mini Code 1.0
- Nemotron 3 Ultra
- MiniMax-M3
- Laguna XS 2.1
- Kimi K3
- Kimi K2.7 Code
- GLM-5.2
- ZAYA1-8B
- Mellum2 12B-A2.5B Thinking
- LFM2.5 8B-A1B
- Gemma 4 12B
- Command A+
- Qwen3.6 35B-A3B
- Qwen3.6 27B
- MiniMax-M2.7
- MiMo-V2.5-Pro
- MiMo-V2.5
- Ling 2.6 1T
- Laguna XS.2
- Kimi K2.6
- Hy3 preview
- Granite 4.1 30B
- GLM-5.1
- DeepSeek-V4-Pro
- DeepSeek-V4-Flash
- Sarvam 30B
- Sarvam 105B
- Nemotron 3 Super
- Nemotron 3 Nano 4B
- LFM2.5 350M
- Gemma 4 E4B
- Gemma 4 E2B
- Gemma 4 31B
- Gemma 4 26B-A4B
- Tiny Aya
- Step 3.5 Flash
- Qwen3.5 397B-A17B
- Nanbeige4.1 3B
- MiniMax-M2.5
- Ling 2.5 1T
- GLM-5
- Trinity Large
- Mistral Small 4
- LongCat-Flash-Lite
- LFM2.5 1.2B
- Kimi K2.5
- Nemotron 3 Nano 30B-A3B
- MiMo-V2-Flash
- GLM-4.7
- DeepSeek-V3.2
- Olmo 3 32B
- Mistral Large 3
- INTELLECT-3
- Ouro 2.6B Thinking
- MiniMax-M2
- Kimi Linear 48B-A3B
- Qwen3-Next 80B-A3B
- Olmo 3 7B
- Grok 2.5
- gpt-oss-20b
- gpt-oss-120b
- Gemma 3 270M
- Qwen3-Coder 30B-A3B
- Kimi K2
- GLM-4.5-Air
- GLM-4.5
- SmolLM3 3B
- Qwen3 8B
- Qwen3 4B
- Qwen3 32B
- Qwen3 30B-A3B
- Qwen3 235B-A22B
- Qwen3 0.6B
- Llama 4 Maverick
- Mistral Small 3.1
- Gemma 3 27B
- MiniMax-Text-01
- DeepSeek-R1
- Phi-4
- OLMo 2 7B
- DeepSeek-V3
- Qwen2.5 72B
- Llama 3.2 3B
- Llama 3.2 1B
- OLMoE-1B-7B
- Llama 3.1 405B
- Gemma 2 27B
- Phi-3-mini
- Mixtral 8x22B
- Llama 3 8B
- DeepSeek-V2
- Grok-1
- Qwen1.5-MoE-A2.7B
- DeepSeekMoE 16B
- Mixtral 8x7B
- Mistral 7B
- Llama 2 70B
- Falcon-7B
- Falcon-40B
- LLaMA 65B
- PaLM 540B
- GPT-NeoX-20B
- GPT-J 6B
Partial RoPE: 60 models in the data set
Rotary embeddings on only part of each head. Newest first; each links to its sourced page.
- Naive-N0.5-Flash
- MiMo-V2.6-Pro-RL
- MiMo-V2.6-Flash-RL
- Ember-1
- DeepSeek-V4.1-Flash
- Qwen3.8-Flash-Next
- Qwen3.8 27B
- Ornith 1.5 35B-A3B
- Ling 3.0 Flash
- Hy4 preview
- GLM-5.3-Flash
- Motif 3 Beta
- Laguna S 2.1
- BTL-3
- MiniMax-M3
- Laguna XS 2.1
- Kimi K3
- Kimi K2.7 Code
- GLM-5.2
- ZAYA1-8B
- Gemma 4 12B
- Qwen3.6 35B-A3B
- Qwen3.6 27B
- MiniMax-M2.7
- MiMo-V2.5-Pro
- MiMo-V2.5
- Ling 2.6 1T
- Laguna XS.2
- Kimi K2.6
- GLM-5.1
- DeepSeek-V4-Pro
- DeepSeek-V4-Flash
- Sarvam 105B
- Gemma 4 E4B
- Gemma 4 E2B
- Gemma 4 31B
- Gemma 4 26B-A4B
- Qwen3.5 397B-A17B
- MiniMax-M2.5
- Ling 2.5 1T
- GLM-5
- Mistral Small 4
- LongCat-Flash-Lite
- Kimi K2.5
- MiMo-V2-Flash
- GLM-4.7
- DeepSeek-V3.2
- Mistral Large 3
- INTELLECT-3
- MiniMax-M2
- Kimi Linear 48B-A3B
- Qwen3-Next 80B-A3B
- Kimi K2
- GLM-4.5-Air
- GLM-4.5
- MiniMax-Text-01
- DeepSeek-R1
- DeepSeek-V3
- DeepSeek-V2
- GPT-NeoX-20B
NoPE layers: 12 models in the data set
Some or all attention layers use no positional encoding. Newest first; each links to its sourced page.
ALiBi: 1 model in the data set
Linear attention biases by distance. Newest first; each links to its sourced page.
Learned absolute: 4 models in the data set
A learned embedding per position. Newest first; each links to its sourced page.
Maths
RoPE writes position into a query (and a key) by rotating each pair , :
Rotations compose, so depends on only. Pair has wavelength ; the interactive counts the pairs with greater than the context length. With partial RoPE, is the rotated fraction of the head, rounded down to an even number.
Code
// src/lib/chapters/model.ts — the wavelength of every rotated pair
const d = rotaryDims(headDim, fraction);
const out: number[] = [];
for (let i = 0; i < d / 2; i++)
out.push(2.0 * Math.PI * theta ** ((2.0 * i) / d));
return out;
The θ values are read from each model's config.json at a pinned Hugging
Face revision (the link under the chart). tests/unit/chapters/model.test.ts
checks the wavelengths against the Python reference to 1e-12 (a power can
round differently in the last bit between maths libraries) and the pair
counts exactly.