llm-architectures-explained
← /learn · 03

Where the norms go

Post-LN, pre-norm, sandwich and post-norm, and QK-norm.

Concept

Every block adds its sub-blocks' outputs (attention, then the feed-forward block) to a residual stream that runs from the embedding to the output head. Normalisation layers keep the numbers on that stream in range. They weigh almost nothing, and they cost almost nothing to run, but where they sit decides whether a deep model trains at all.

  • Post-LN, the original Transformer: normalise after each residual addition, x←LN(x+f(x))x \leftarrow \mathrm{LN}(x + f(x)). Every norm rescales the whole stream, so whatever the embedding put there is shrunk again at each of the dozens of norms above it. Xiong et al. (2020) showed that at initialisation the gradients near the output are large, which is why Post-LN needs a learning-rate warm-up.
  • Pre-norm, the placement of most open models in this data set: normalise the sub-block's input, x←x+f(LN(x))x \leftarrow x + f(\mathrm{LN}(x)). The stream itself is never normalised, so the embedding passes straight through, and the same paper showed the gradients are well behaved at initialisation, so the warm-up can go. The price: the stream's size grows with depth, so each later layer's update is a smaller fraction of it, and if a sub-block's output grows during training, nothing caps it.
  • A norm on the sub-block's output caps it: sandwich norm (Gemma 2 and 3: RMSNorm on both the input and the output of each sub-layer) and OLMo 2's reordered norm (x←x+RMSNorm(f(x))x \leftarrow x + \mathrm{RMSNorm}(f(x)), outputs only), a strategy first proposed to stabilise training, as the OLMo 2 report notes.
  • QK-norm normalises queries and keys inside attention, before their dot product, which "avoids attention logits being too large, which can lead to training loss divergence" (OLMo 2, p. 5). Gemma 3 replaced Gemma 2's logit soft-capping with it.
  • Parallel blocks (chapter 8) feed one norm's output to attention and the feed-forward block side by side.

The interactive runs the textbook variance argument behind those statements: a unit-size embedding, and every sub-block adding an output independent of the stream, whose size is a gain times its input's. At 32 layers and gain 1, pre-norm's stream ends 8.06 times the embedding's size, the last sub-block changes it by 12.5%, and the embedding is still 12.4% of it. Post-LN keeps the stream at 1, but the embedding's share at the top is 2−322^{-32}. Raise the gain to 2 and pre-norm's stream doubles to 16.03, while the output-norm stream does not move.

Loading the interactive…

Pre-norm: 109 models in the data set

A norm before each sub-block. Newest first; each links to its sourced page.

Sandwich norm: 11 models in the data set

Norms both before and after each sub-block. Newest first; each links to its sourced page.

Post-norm: 3 models in the data set

Norms after each sub-block, inside the residual (OLMo 2). Newest first; each links to its sourced page.

QK-norm: 48 models in the data set

Queries and keys normalised before the dot product. Newest first; each links to its sourced page.

Post-LN: 2 models in the data set

The original Transformer: a norm after each residual addition. Newest first; each links to its sourced page.

Maths

Let vkv_k be the stream's variance after kk sub-blocks, v0=1v_0 = 1, and let a sub-block add an output independent of the stream with variance g2g^2 times its input's (γ2\gamma^2 when its output is normalised and scaled by γ\gamma). Then, over K=2lK = 2l sub-blocks:

Placementvkv_kupdate / streamembedding share at the top
Post-LN11gg(1+g2)−K/2(1 + g^2)^{-K/2}
Pre-norm1+kg21 + k g^2g/1+(k−1)g2g / \sqrt{1 + (k-1)g^2}(1+Kg2)−1/2(1 + K g^2)^{-1/2}
Output norm1+kγ21 + k \gamma^2γ/1+(k−1)γ2\gamma / \sqrt{1 + (k-1)\gamma^2}(1+Kγ2)−1/2(1 + K\gamma^2)^{-1/2}

The output-norm row does not contain gg: that is the whole argument for it. (In this idealisation sandwich norm and OLMo 2's output-only norm are the same row; they differ in whether the sub-block's input is normalised too.)

Code

// src/lib/chapters/model.ts — one sub-block of the profile
if (placement === "post-ln") {
  update.push(gain / Math.sqrt(v));
  const total = v + gain * gain;
  emb = emb / Math.sqrt(total);
  v = 1.0;
} else {
  const add = placement === "pre-norm" ? gain : gamma;
  update.push(add / Math.sqrt(v));
  v = v + add * add;
  emb = 1.0 / Math.sqrt(v);
}