llm-architectures-explained
← /learn · 05

Depth and width

How layer count and model width trade parameters, KV and FLOPs.

Concept

Two numbers set a decoder's size: its depth (layers) and its width (the model dimension dd). A layer's weights grow as d2d^2, so a fixed parameter budget can be spent on a few wide layers or many narrow ones. The data set spans both ends: Llama 3.1 405B stacks 126 layers of width 16,384, a width 130 times its depth; Gemma 3 270M stacks 18 layers of width 640, about 36 times.

For quality, the shape matters less than you might expect, within limits. Kaplan et al. (2020) found that "architectural details such as network width or depth have minimal effects within a wide range"; below a billion parameters, MobileLLM (Liu et al., 2024) argued that architecture matters more, and built its models deep and thin.

For cost, the shape matters in ways the parameter count hides. The interactive builds Llama-style decoders of any shape and costs them:

  • The KV cache grows with depth, not width. Each layer caches its keys and values, and with the number of KV heads fixed (8 of 128 dimensions here) a layer's cache per token does not depend on dd. An 80-layer, 2,048-wide model caches 320 KiB per token; a 12-layer, 8,192-wide one 48 KiB, though it has more than twice the parameters (12.4B against 5.06B).
  • Attention's share of decode grows with depth too. At a 32K context the deep, thin model spends 69.1% of its decode FLOPs in attention, the Llama 3 8B shape 53.4%, the shallow, wide model 36.3%.
  • The embedding tables grow with width only. With a 128,256-token vocabulary they are 10.4% of the deep, thin model's parameters and 17.0% of the shallow, wide one's (untied input and output tables).
  • Depth costs latency. Layers run one after another for every token, so at batch 1 a deeper model has more sequential steps per token; width can be split across GPUs (tensor parallelism), depth only pipelined.

Loading the interactive…

Maths

For a Llama-style layer of width dd with nh=d/128n_h = d/128 query heads, nkvn_{kv} KV heads of dh=128d_h = 128 and a gated FFN of width dffd_{ff}:

Player=2d2+2d nkvdh⏟attention+3d dff⏟FFN,KV per token=2 l nkv dh⋅b,P_{\text{layer}} = \underbrace{2 d^2 + 2 d\, n_{kv} d_h}_{\text{attention}} + \underbrace{3 d\, d_{ff}}_{\text{FFN}}, \qquad \text{KV per token} = 2\, l\, n_{kv}\, d_h \cdot b,

with ll layers and bb bytes per value. Parameters scale as ld2l d^2, the cache as ll alone. A decode step costs 2Pactive2P_{\text{active}} FLOPs for the matrix multiplications plus 4nhdh t=4d t4 n_h d_h\, t = 4 d\, t per layer at context tt, so the attention share is roughly 4ldt2P+4ldt\dfrac{4 l d t}{2 P + 4 l d t}, which grows with ll at a fixed PP.

Code

// src/lib/chapters/model.ts — the shape the sliders build (excerpt)
  const heads = Math.floor(dModel / 128);
  return {
    kind: "decoder",
    d_model: dModel,
    vocab,
    tied_embeddings: false,