/learn
Architectures, axis by axis
One chapter per way LLM architectures differ. Each has a live interactive driven by this site's cost model (checked against its Python reference) and links to the models in the data set that use the feature. Toggle layers (Concept / Maths / Code) inside any chapter to choose how deep to go. New to transformers? Start with the Transformer Decoder Explainer; for what happens when a model is served, read LLM Inference Explained.
MHA, GQA, MQA, MLA, sliding windows, sparse attention, and linear and state-space hybrids.
Learned, sinusoidal, ALiBi, RoPE, partial RoPE and NoPE layers.
Post-LN, pre-norm, sandwich and post-norm, and QK-norm.
Expert counts, active experts, shared experts and dense first layers.
How layer count and model width trade parameters, KV and FLOPs.
RoPE scaling, windows, compressed and sparse attention, and hybrids, at 1M tokens.
Extra heads that draft the next tokens, and what they buy.
Reusing the layer stack, and running attention and FFN side by side.
T5's split, and DeepSeek-V4.1-Flash's encoder-only prefill, simulated live.