Concept
Every block adds its sub-blocks' outputs (attention, then the feed-forward block) to a residual stream that runs from the embedding to the output head. Normalisation layers keep the numbers on that stream in range. They weigh almost nothing, and they cost almost nothing to run, but where they sit decides whether a deep model trains at all.
- Post-LN, the original Transformer: normalise after each residual addition, . Every norm rescales the whole stream, so whatever the embedding put there is shrunk again at each of the dozens of norms above it. Xiong et al. (2020) showed that at initialisation the gradients near the output are large, which is why Post-LN needs a learning-rate warm-up.
- Pre-norm, the placement of most open models in this data set: normalise the sub-block's input, . The stream itself is never normalised, so the embedding passes straight through, and the same paper showed the gradients are well behaved at initialisation, so the warm-up can go. The price: the stream's size grows with depth, so each later layer's update is a smaller fraction of it, and if a sub-block's output grows during training, nothing caps it.
- A norm on the sub-block's output caps it: sandwich norm (Gemma 2 and 3: RMSNorm on both the input and the output of each sub-layer) and OLMo 2's reordered norm (, outputs only), a strategy first proposed to stabilise training, as the OLMo 2 report notes.
- QK-norm normalises queries and keys inside attention, before their dot product, which "avoids attention logits being too large, which can lead to training loss divergence" (OLMo 2, p. 5). Gemma 3 replaced Gemma 2's logit soft-capping with it.
- Parallel blocks (chapter 8) feed one norm's output to attention and the feed-forward block side by side.
The interactive runs the textbook variance argument behind those statements: a unit-size embedding, and every sub-block adding an output independent of the stream, whose size is a gain times its input's. At 32 layers and gain 1, pre-norm's stream ends 8.06 times the embedding's size, the last sub-block changes it by 12.5%, and the embedding is still 12.4% of it. Post-LN keeps the stream at 1, but the embedding's share at the top is . Raise the gain to 2 and pre-norm's stream doubles to 16.03, while the output-norm stream does not move.
Loading the interactive…
Pre-norm: 109 models in the data set
A norm before each sub-block. Newest first; each links to its sourced page.
- Naive-N0.5-Flash
- MiMo-V2.6-Pro-RL
- MiMo-V2.6-Flash-RL
- Ember-1
- Qwen3.8 27B
- Ornith 1.5 35B-A3B
- Nemotron 3.5 Lightning
- Ling 3.0 Flash
- Hy4 preview
- GLM-5.3-Flash
- Soofi S
- Solar Open 2
- Nanbeige4.2 3B
- Motif 3 Beta
- Laguna S 2.1
- BTL-3
- Antares 1B
- VibeThinker-3B
- Nemotron 3 Ultra
- Laguna XS 2.1
- Kimi K3
- Kimi K2.7 Code
- GLM-5.2
- ZAYA1-8B
- Mellum2 12B-A2.5B Thinking
- LFM2.5 8B-A1B
- Qwen3.6 35B-A3B
- Qwen3.6 27B
- MiniMax-M2.7
- MiMo-V2.5-Pro
- MiMo-V2.5
- Ling 2.6 1T
- Laguna XS.2
- Kimi K2.6
- Hy3 preview
- Granite 4.1 30B
- GLM-5.1
- DeepSeek-V4-Pro
- DeepSeek-V4-Flash
- Sarvam 30B
- Sarvam 105B
- Nemotron 3 Super
- Nemotron 3 Nano 4B
- LFM2.5 350M
- Step 3.5 Flash
- Qwen3.5 397B-A17B
- Nanbeige4.1 3B
- MiniMax-M2.5
- Ling 2.5 1T
- GLM-5
- Mistral Small 4
- LFM2.5 1.2B
- Kimi K2.5
- Nemotron 3 Nano 30B-A3B
- MiMo-V2-Flash
- GLM-4.7
- DeepSeek-V3.2
- INTELLECT-3
- MiniMax-M2
- Kimi Linear 48B-A3B
- Qwen3-Next 80B-A3B
- gpt-oss-20b
- gpt-oss-120b
- Qwen3-Coder 30B-A3B
- Kimi K2
- GLM-4.5-Air
- GLM-4.5
- SmolLM3 3B
- Qwen3 8B
- Qwen3 4B
- Qwen3 32B
- Qwen3 30B-A3B
- Qwen3 235B-A22B
- Qwen3 0.6B
- Llama 4 Maverick
- Mistral Small 3.1
- MiniMax-Text-01
- DeepSeek-R1
- xLSTM 7B
- Phi-4
- DeepSeek-V3
- Qwen2.5 72B
- Llama 3.2 3B
- Llama 3.2 1B
- RWKV-6 Finch 14B
- OLMoE-1B-7B
- Llama 3.1 405B
- Mamba-2 2.7B
- Phi-3-mini
- Mixtral 8x22B
- Llama 3 8B
- DeepSeek-V2
- Jamba v0.1
- Qwen1.5-MoE-A2.7B
- DeepSeekMoE 16B
- Mixtral 8x7B
- Mamba 2.8B
- Mistral 7B
- Llama 2 70B
- RWKV-4 14B
- LLaMA 65B
- BLOOM
- OPT-66B
- Chinchilla 70B
- GLaM (64B/64E)
- Switch-C
- GPT-3 175B
- T5 11B
- GPT-2 XL
Sandwich norm: 11 models in the data set
Norms both before and after each sub-block. Newest first; each links to its sourced page.
Post-norm: 3 models in the data set
Norms after each sub-block, inside the residual (OLMo 2). Newest first; each links to its sourced page.
QK-norm: 48 models in the data set
Queries and keys normalised before the dot product. Newest first; each links to its sourced page.
- Qwen3.8-Flash-Next
- Qwen3.8 27B
- Ornith 1.5 35B-A3B
- Ling 3.0 Flash
- Nanbeige4.2 3B
- Laguna S 2.1
- BTL-3
- MiniMax-M3
- Laguna XS 2.1
- Mellum2 12B-A2.5B Thinking
- LFM2.5 8B-A1B
- Gemma 4 12B
- Qwen3.6 35B-A3B
- Qwen3.6 27B
- MiniMax-M2.7
- Ling 2.6 1T
- Laguna XS.2
- Hy3 preview
- Sarvam 30B
- Sarvam 105B
- LFM2.5 350M
- Gemma 4 E4B
- Gemma 4 E2B
- Gemma 4 31B
- Gemma 4 26B-A4B
- Step 3.5 Flash
- Qwen3.5 397B-A17B
- MiniMax-M2.5
- Ling 2.5 1T
- Trinity Large
- LFM2.5 1.2B
- GLM-4.7
- Olmo 3 32B
- MiniMax-M2
- Qwen3-Next 80B-A3B
- Olmo 3 7B
- Gemma 3 270M
- Qwen3-Coder 30B-A3B
- GLM-4.5
- Qwen3 8B
- Qwen3 4B
- Qwen3 32B
- Qwen3 30B-A3B
- Qwen3 235B-A22B
- Qwen3 0.6B
- Gemma 3 27B
- OLMo 2 7B
- OLMoE-1B-7B
Post-LN: 2 models in the data set
The original Transformer: a norm after each residual addition. Newest first; each links to its sourced page.
Maths
Let be the stream's variance after sub-blocks, , and let a sub-block add an output independent of the stream with variance times its input's ( when its output is normalised and scaled by ). Then, over sub-blocks:
| Placement | update / stream | embedding share at the top | |
|---|---|---|---|
| Post-LN | |||
| Pre-norm | |||
| Output norm |
The output-norm row does not contain : that is the whole argument for it. (In this idealisation sandwich norm and OLMo 2's output-only norm are the same row; they differ in whether the sub-block's input is normalised too.)
Code
// src/lib/chapters/model.ts — one sub-block of the profile
if (placement === "post-ln") {
update.push(gain / Math.sqrt(v));
const total = v + gain * gain;
emb = emb / Math.sqrt(total);
v = 1.0;
} else {
const add = placement === "pre-norm" ? gain : gamma;
update.push(add / Math.sqrt(v));
v = v + add * add;
emb = 1.0 / Math.sqrt(v);
}