llm-architectures-explained

/models

Transformer (base)

Google · Transformer · open weights · encoder-decoder

Facts and where they come from

Released2017-06paperpaperarXiv v1, June 2017
Licencenot disclosednot disclosed
Total parameters65MpaperpaperTable 3: base model, 65 x 10^6 parameters
Active parametersnot disclosednot disclosed
Context lengthnot disclosednot disclosed
Norm placementpost-lncodemodelling codeAttention Is All You Need §3.1: LayerNorm(x + Sublayer(x))
Norm typeLayerNormcodemodelling codeAttention Is All You Need §3.1: LayerNorm(x + Sublayer(x))
QK-normnocodemodelling codeno q/k normalisation in the attention block
Positional encodingsinusoidalpaperpaper§3.5: sinusoidal positional encodings
Parallel attention and MLPnocodemodelling codeAttention Is All You Need §3.1: LayerNorm(x + Sublayer(x))

Architecture, drawn from the data

MHA 8q/8kv. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

Transformer (base): layer stack and blocksencoder (6)mixer / FFNlayer 0: MHA: 8 query / 8 KV heads · head 64layer 0: MLP: 2048layer 1: MHA: 8 query / 8 KV heads · head 64layer 1: MLP: 2048layer 2: MHA: 8 query / 8 KV heads · head 64layer 2: MLP: 2048layer 3: MHA: 8 query / 8 KV heads · head 64layer 3: MLP: 2048layer 4: MHA: 8 query / 8 KV heads · head 64layer 4: MLP: 2048layer 5: MHA: 8 query / 8 KV heads · head 64layer 5: MLP: 2048035decoder (6)mixer / FFNlayer 0: MHA: 8 query / 8 KV heads · head 64layer 0: MLP: 2048layer 1: MHA: 8 query / 8 KV heads · head 64layer 1: MLP: 2048layer 2: MHA: 8 query / 8 KV heads · head 64layer 2: MLP: 2048layer 3: MHA: 8 query / 8 KV heads · head 64layer 3: MLP: 2048layer 4: MHA: 8 query / 8 KV heads · head 64layer 4: MLP: 2048layer 5: MHA: 8 query / 8 KV heads · head 64layer 5: MLP: 2048035× 6MHA: 8 query / 8 KV heads · head 64norm+MLP: 2048norm+full attentiondense FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)63M
Active per token (modelled)63M
Without embeddings and output head44.1M total, 44.1M active
KV cache per token, BF16 (layers that grow with context)24 KiB
KV cache + state at 128K tokens, BF163 GiB
Decode FLOPs per token at 4K context227 MFLOP
Prefill FLOPs for a 4K prompt515 GFLOP

KV cache against context

Transformer (base): KV cache bytes against context length101001,00010,000100,00098 KiB980 KiB9.5 MiB95 MiB950 MiBcontext (tokens)KV cache + state (BF16)Transformer (base)

Compare with other models →

Every architecture field

FieldValueSource
d_model512paperpaperTable 3 (base): d_model 512
vocab37,000paperpaper§5.1: shared source-target vocabulary of about 37000 tokens (EN-DE)
tied_embeddingstruecodemodelling codetransformers 5.18.0 t5: tie_word_embeddings default true
mixers.full.typeattncodemodelling codeattention
mixers.full.heads8paperpaperTable 3 (base): h 8
mixers.full.kv_heads8paperpaperTable 3 (base): h 8
mixers.full.head_dim64paperpaperTable 3 (base): d_k = d_v = 64
ffns.dense.typedensecodemodelling codeMLP
ffns.dense.d_ff2,048paperpaperTable 3 (base): d_ff 2048
ffns.dense.gatedfalsecodemodelling codefeed_forward_proj
layout6× full/densepaperpaperTable 3 (base): N = 6
encoder_layout6× full/densepaperpaperTable 3 (base): N = 6 (encoder and decoder)
norms_per_layer2codemodelling codepre-norm RMSNorm

Sources