Concept
Two numbers set a decoder's size: its depth (layers) and its width (the model dimension ). A layer's weights grow as , so a fixed parameter budget can be spent on a few wide layers or many narrow ones. The data set spans both ends: Llama 3.1 405B stacks 126 layers of width 16,384, a width 130 times its depth; Gemma 3 270M stacks 18 layers of width 640, about 36 times.
For quality, the shape matters less than you might expect, within limits. Kaplan et al. (2020) found that "architectural details such as network width or depth have minimal effects within a wide range"; below a billion parameters, MobileLLM (Liu et al., 2024) argued that architecture matters more, and built its models deep and thin.
For cost, the shape matters in ways the parameter count hides. The interactive builds Llama-style decoders of any shape and costs them:
- The KV cache grows with depth, not width. Each layer caches its keys and values, and with the number of KV heads fixed (8 of 128 dimensions here) a layer's cache per token does not depend on . An 80-layer, 2,048-wide model caches 320 KiB per token; a 12-layer, 8,192-wide one 48 KiB, though it has more than twice the parameters (12.4B against 5.06B).
- Attention's share of decode grows with depth too. At a 32K context the deep, thin model spends 69.1% of its decode FLOPs in attention, the Llama 3 8B shape 53.4%, the shallow, wide model 36.3%.
- The embedding tables grow with width only. With a 128,256-token vocabulary they are 10.4% of the deep, thin model's parameters and 17.0% of the shallow, wide one's (untied input and output tables).
- Depth costs latency. Layers run one after another for every token, so at batch 1 a deeper model has more sequential steps per token; width can be split across GPUs (tensor parallelism), depth only pipelined.
Loading the interactive…
Maths
For a Llama-style layer of width with query heads, KV heads of and a gated FFN of width :
with layers and bytes per value. Parameters scale as , the cache as alone. A decode step costs FLOPs for the matrix multiplications plus per layer at context , so the attention share is roughly , which grows with at a fixed .
Code
// src/lib/chapters/model.ts — the shape the sliders build (excerpt)
const heads = Math.floor(dModel / 128);
return {
kind: "decoder",
d_model: dModel,
vocab,
tied_embeddings: false,