LLM Architectures Explained
How language-model architectures differ, model by model.
Every large language model is a stack of the same few parts, chosen differently: how attention shares its keys and values, how positions are encoded, where the norms sit, whether the feed-forward block is one network or many experts, how deep and how wide. This site records those choices for 160 models, each value traced to the model’s own configuration, paper or announcement, and computes what they cost.
- models
- 160
- gallery checklist names covered
- 109/109
- configs pinned to a Hugging Face commit
- 122
- closed models, sizes not disclosed unless stated
- 23
Nine ways models differ
One chapter per axis of variation, from attention to the causal encoder-decoder, each with a live interactive and the models that use it.
Start learning →
Every model
160 models in one table: size, depth, width, attention, KV cache per token. Filter by any design choice.
Open the table →
Compare and calculate
Pick up to four models: diagrams drawn from their data, KV cache against context, prefill and decode cost.
Compare models →
How the variations spread
From the 2017 Transformer to this year: when GQA, MoE, MLA, sliding windows and linear attention took hold.
See the timeline →
Part of a family of companion sites: the Transformer Decoder Explainer (one forward pass), LLM Inference Explained (serving it), GPU Kernels Explained (how a GPU executes it), Numerics Explained (the number formats it runs in) and Systolic Arrays Explained (the matrix hardware of TPUs). How this site was built, and how to check its data: about.