/timeline
How the variations spread
Each row is one design choice; each dot is a model in this data set that makes it, placed by release month (for open models, the month the Hugging Face repository was created unless a paper or announcement date is recorded). Hover or focus a dot for the model; follow it for the sources. Filled dots are open weights, hollow dots closed.
Earliest in this data set
The first model here with each feature. This is a property of the data set, not a claim about who invented it: many ideas appeared first in papers or in models not listed.
- Encoder-decoder: Transformer (base) (2017-06)
- Learned absolute: BERT-Large (2018-10)
- RoPE: GPT-J 6B (2021-06)
- Partial RoPE: GPT-NeoX-20B (2022-04)
- NoPE layers: Jamba v0.1 (2024-03)
- Parallel block: GPT-J 6B (2021-06)
- GQA: Falcon-40B (2023-05)
- MQA: PaLM 540B (2022-04)
- Sliding window: Mistral 7B (2023-09)
- Chunked attention: Llama 4 Maverick (2025-04)
- MoE: Switch-C (2021-01)
- Shared expert: DeepSeekMoE 16B (2024-01)
- MLA: DeepSeek-V2 (2024-04)
- Mamba (SSM): Mamba 2.8B (2023-12)
- Linear attention: RWKV-4 14B (2023-05)
- DeltaNet / KDA: Qwen3-Next 80B-A3B (2025-09)
- Hybrid: Jamba v0.1 (2024-03)
- Sparse attention: DeepSeek-V3.2 (2025-12)
- QK-norm: OLMoE-1B-7B (2024-07)
- Post-norm: OLMo 2 7B (2024-12)
- Sandwich norm: Gemma 2 27B (2024-06)
- MTP: DeepSeek-V3 (2024-12)
- Looped: Ouro 2.6B Thinking (2025-10)
- Causal encoder-decoder: DeepSeek-V4.1-Flash (2026-09)