llm-architectures-explained

/models

BERT-Large

Google · BERT · open weights · encoder

Facts and where they come from

Released2018-10paperarXiv 1810.04805arXiv v1, October 2018
Licenceapache-2.0config.jsonconfig.jsonREADME metadata: license
Total parameters340MpaperarXiv 1810.04805§3: BERT-LARGE (L=24, H=1024, A=16, total parameters=340M)
Active parametersnot disclosednot disclosed
Context length512 tokensconfig.jsonconfig.jsonmax_position_embeddings
Norm placementpost-lncodemodelling codetransformers 5.18.0 bert: LayerNorm after each residual add
Norm typeLayerNormcodemodelling codetransformers 5.18.0 bert: LayerNorm after each residual add
QK-normnocodemodelling codeno q/k normalisation in the attention block
Positional encodinglearned absolutecodemodelling codetransformers 5.18.0 bert: learned absolute position embeddings
Parallel attention and MLPnocodemodelling codetransformers 5.18.0 bert: LayerNorm after each residual add

Architecture, drawn from the data

MHA 16q/16kv. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

BERT-Large: layer stack and blockslayers (24)mixer / FFNlayer 0: MHA: 16 query / 16 KV heads · head 64layer 0: MLP: 4096layer 1: MHA: 16 query / 16 KV heads · head 64layer 1: MLP: 4096layer 2: MHA: 16 query / 16 KV heads · head 64layer 2: MLP: 4096layer 3: MHA: 16 query / 16 KV heads · head 64layer 3: MLP: 4096layer 4: MHA: 16 query / 16 KV heads · head 64layer 4: MLP: 4096layer 5: MHA: 16 query / 16 KV heads · head 64layer 5: MLP: 4096layer 6: MHA: 16 query / 16 KV heads · head 64layer 6: MLP: 4096layer 7: MHA: 16 query / 16 KV heads · head 64layer 7: MLP: 4096layer 8: MHA: 16 query / 16 KV heads · head 64layer 8: MLP: 4096layer 9: MHA: 16 query / 16 KV heads · head 64layer 9: MLP: 4096layer 10: MHA: 16 query / 16 KV heads · head 64layer 10: MLP: 4096layer 11: MHA: 16 query / 16 KV heads · head 64layer 11: MLP: 4096layer 12: MHA: 16 query / 16 KV heads · head 64layer 12: MLP: 4096layer 13: MHA: 16 query / 16 KV heads · head 64layer 13: MLP: 4096layer 14: MHA: 16 query / 16 KV heads · head 64layer 14: MLP: 4096layer 15: MHA: 16 query / 16 KV heads · head 64layer 15: MLP: 4096layer 16: MHA: 16 query / 16 KV heads · head 64layer 16: MLP: 4096layer 17: MHA: 16 query / 16 KV heads · head 64layer 17: MLP: 4096layer 18: MHA: 16 query / 16 KV heads · head 64layer 18: MLP: 4096layer 19: MHA: 16 query / 16 KV heads · head 64layer 19: MLP: 4096layer 20: MHA: 16 query / 16 KV heads · head 64layer 20: MLP: 4096layer 21: MHA: 16 query / 16 KV heads · head 64layer 21: MLP: 4096layer 22: MHA: 16 query / 16 KV heads · head 64layer 22: MLP: 4096layer 23: MHA: 16 query / 16 KV heads · head 64layer 23: MLP: 409601223× 24MHA: 16 query / 16 KV heads · head 64norm+MLP: 4096norm+full attentiondense FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)334M
Active per token (modelled)334M
Without embeddings and output head302M total, 302M active
Published weights (Hugging Face count)336M
KV cache per token, BF16 (layers that grow with context)96 KiB
KV cache + state at 512 tokens, BF1648 MiB
Decode FLOPs per token at 4K context1.07 GFLOP
Prefill FLOPs for a 4K prompt3.3 TFLOP

KV cache against context

BERT-Large: KV cache bytes against context length101001,000980 KiB9.5 MiB95 MiBcontext (tokens)KV cache + state (BF16)BERT-Large

Compare with other models →

Every architecture field

FieldValueSource
d_model1,024config.jsonconfig.jsonhidden_size
vocab30,522config.jsonconfig.jsonvocab_size
tied_embeddingstruecodemodelling codeMLM head tied to word embeddings
mixers.full.typeattncodemodelling codeattention block
mixers.full.heads16config.jsonconfig.jsonnum_attention_heads
mixers.full.kv_heads16config.jsonconfig.jsonnum_attention_heads
mixers.full.head_dim64codemodelling codetransformers 5.18.0: head_dim = hidden_size / num_attention_heads
mixers.full.biastruecodemodelling codebiases
ffns.dense.typedensecodemodelling codeMLP block
ffns.dense.d_ff4,096config.jsonconfig.jsonintermediate_size
ffns.dense.gatedfalsecodemodelling codeMLP: two matrices, no gate
ffns.dense.biastruecodemodelling codeMLP has biases
layout24× full/denseconfig.jsonconfig.jsonnum_hidden_layers
extra_embedding_params526,336codemodelling codeposition and token-type embeddings

Sources