/about
About this site
LLM Architectures Explained records how 160 language models are built and what that costs. It is the third of a family of companion sites: the Transformer Decoder Explainer shows one forward pass, LLM Inference Explained shows how a model is served, GPU Kernels Explained shows how a GPU executes it, Numerics Explained shows the number formats it runs in, Systolic Arrays Explained shows the matrix hardware of TPUs, and Inference Trade-offs Explained measures which serving lever helps which metric. This one shows how the models themselves differ.
Where every value comes from
Every value carries a status and a link to its source, at the place it was read:
- config.json read from the model’s
config.jsonat a pinned Hugging Face commit (122 models). A test re-reads each such value from the stored snapshot. - lab stated by the lab: its model card at the same pinned commit, its GitHub, its documentation or its announcement.
- paper from the model’s paper, by section or table. Every arXiv identifier was checked against arXiv.
- code a rule from the model’s modelling code (a default the configuration leaves out, or where the norms sit), citing the file.
- not disclosed nothing credible is published.
- reported estimate a third-party estimate for a closed model; see below.
Where a configuration is gated on Hugging Face, the values come instead from the lab’s own GitHub (Meta’sllama-models, Google DeepMind’s gemma, xAI’s grok-1) or from the paper, transcribed into a file that names the line or table for every value. The data, the snapshots and the transcriptions are all in the repository, with a JSON Schema that CI validates.
The gallery checklist
Sebastian Raschka’s LLM Architecture Gallery is good related reading. This site uses it only as a checklist: its 109 model names are stored, and a test checks that each one maps to a model here. No figure, diagram or text is taken from it; every diagram here is generated from the data, and every fact comes from the model’s own sources. The other 51 models add the history the gallery skips (the 2017 Transformer, BERT, T5, GPT-3, Switch, GLaM, PaLM, Chinchilla, BLOOM, Llama, Mistral and Mixtral, Mamba, RWKV, Jamba and more) and closed frontier models.
Closed models and reported estimates
Closed models are listed with what their labs disclose: usually the context window and the release date, and sometimes that the model is a mixture of experts. Where a third party has published a size estimate, it is shown as a reported estimate: never as a fact, always with its source, its date and the confidence its source gives. There are 6 such values, all from one source whose authors say they cannot vouch for them. The calculator ignores them unless you switch them on. Where nothing credible exists, the site says “not disclosed”.
The cost model
For each model with published dimensions (142 of 160), the cost model computes parameter counts, KV-cache and recurrent-state bytes, prefill and decode FLOPs, and the bytes a decode step reads. It is a Python reference (reference/arch_model.py) with unit tests for every closed form, and a TypeScript port that this site runs; CI checks that the port reproduces every number in the reference’s fixtures exactly, for every model. Conventions:
- a matrix-vector product of an m-vector with an m × n matrix costs 2mn FLOPs; norms, activations, softmax and routing are not counted;
- attention is counted in its expanded form for every type (MLA’s absorbed decode trades these FLOPs for latent-space ones);
- sliding-window layers keep only their window; chunked layers (Llama 4) keep one chunk, and each query reads only its own chunk so far; sparse attention reads only the selected entries, plus its indexer keys in full;
- linear attention, Mamba and short convolutions keep a fixed-size state, not a cache;
- a causal encoder-decoder (DeepSeek-V4.1-Flash) runs only its encoder half over the prompt, projects the decoder’s keys and values from the encoder output, and replays the last tokens through the decoder half;
- batch size 1; serving at scale changes the balance between weights and cache (see LLM Inference Explained).
How the counts are checked
The modelled parameter count of 80 of the 86 open text models whose published weights can be counted (unpacked, no vision tower) is within 0.5% of the number of parameters in those weights, at the pinned commit. Every stated total is also checked against the model (within 5%, or between the counts with and without embeddings, since labs differ on that), with each exception named in the tests. Some parts are approximated and say so on their model pages: compressed convolutional attention (ZAYA1), Kimi Delta Attention’s gate projections, n-gram embedding tables, GDLA (Motif) and the weights of multi-token-prediction layers.
The chapters and the simulator
The nine chapters each drive an interactive with this cost model, or with a closed form from the same Python reference (reference/chapter_model.py, ported line by line and checked against its fixtures). The encoder-decoder chapter also runs Disaggregated_Inference_Sim’s own JavaScript engine, copied byte for byte from commit e674e18 (the one that added the causal encoder-decoder option), with parity tests against the Python package at that commit, and reruns its capacity search in the browser, rate for rate. Its numbers are for an illustrative dense 70B-shaped proxy, not for any lab’s model.
Freshness
A weekly workflow re-fetches every pinned configuration and reports any model whose repository has moved on, so the data can be re-pinned deliberately rather than drift.
How it is built
Next.js 14 with strict TypeScript and Tailwind, the design system of the companion sites, and diagrams and charts in plain SVG generated from the data. There is no database and no sign-in. Vitest checks the cost model and the data helpers, pytest the reference and the data, and Playwright every page at desktop and phone widths in light and dark mode. Start at the models.