llm-architectures-explained

/models

Falcon-40B

TII · Falcon · open weights

Facts and where they come from

Released2023-05config.jsonconfig.jsonHugging Face repository creation date (api.createdAt)
Licenceapache-2.0config.jsonconfig.jsonREADME metadata: license
Total parameters40Blabmodel cardREADME: Falcon-40B is a 40B parameters causal decoder-only model
Active parametersnot disclosednot disclosed
Context lengthnot disclosednot disclosed
Norm placementparallelcodemodelling coderepo modeling_falcon.py: parallel attention and MLP
Norm typeLayerNormcodemodelling coderepo modeling_falcon.py: parallel attention and MLP
QK-normnocodemodelling codeno q/k normalisation in the attention block
Positional encodingRoPEcodemodelling coderotary on the full head (default)
Parallel attention and MLPyescodemodelling coderepo modeling_falcon.py: parallel attention and MLP

Architecture, drawn from the data

GQA 128q/8kv. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

Falcon-40B: layer stack and blockslayers (60)mixer / FFNlayer 0: GQA: 128 query / 8 KV heads · head 64layer 0: MLP: 32768layer 1: GQA: 128 query / 8 KV heads · head 64layer 1: MLP: 32768layer 2: GQA: 128 query / 8 KV heads · head 64layer 2: MLP: 32768layer 3: GQA: 128 query / 8 KV heads · head 64layer 3: MLP: 32768layer 4: GQA: 128 query / 8 KV heads · head 64layer 4: MLP: 32768layer 5: GQA: 128 query / 8 KV heads · head 64layer 5: MLP: 32768layer 6: GQA: 128 query / 8 KV heads · head 64layer 6: MLP: 32768layer 7: GQA: 128 query / 8 KV heads · head 64layer 7: MLP: 32768layer 8: GQA: 128 query / 8 KV heads · head 64layer 8: MLP: 32768layer 9: GQA: 128 query / 8 KV heads · head 64layer 9: MLP: 32768layer 10: GQA: 128 query / 8 KV heads · head 64layer 10: MLP: 32768layer 11: GQA: 128 query / 8 KV heads · head 64layer 11: MLP: 32768layer 12: GQA: 128 query / 8 KV heads · head 64layer 12: MLP: 32768layer 13: GQA: 128 query / 8 KV heads · head 64layer 13: MLP: 32768layer 14: GQA: 128 query / 8 KV heads · head 64layer 14: MLP: 32768layer 15: GQA: 128 query / 8 KV heads · head 64layer 15: MLP: 32768layer 16: GQA: 128 query / 8 KV heads · head 64layer 16: MLP: 32768layer 17: GQA: 128 query / 8 KV heads · head 64layer 17: MLP: 32768layer 18: GQA: 128 query / 8 KV heads · head 64layer 18: MLP: 32768layer 19: GQA: 128 query / 8 KV heads · head 64layer 19: MLP: 32768layer 20: GQA: 128 query / 8 KV heads · head 64layer 20: MLP: 32768layer 21: GQA: 128 query / 8 KV heads · head 64layer 21: MLP: 32768layer 22: GQA: 128 query / 8 KV heads · head 64layer 22: MLP: 32768layer 23: GQA: 128 query / 8 KV heads · head 64layer 23: MLP: 32768layer 24: GQA: 128 query / 8 KV heads · head 64layer 24: MLP: 32768layer 25: GQA: 128 query / 8 KV heads · head 64layer 25: MLP: 32768layer 26: GQA: 128 query / 8 KV heads · head 64layer 26: MLP: 32768layer 27: GQA: 128 query / 8 KV heads · head 64layer 27: MLP: 32768layer 28: GQA: 128 query / 8 KV heads · head 64layer 28: MLP: 32768layer 29: GQA: 128 query / 8 KV heads · head 64layer 29: MLP: 32768layer 30: GQA: 128 query / 8 KV heads · head 64layer 30: MLP: 32768layer 31: GQA: 128 query / 8 KV heads · head 64layer 31: MLP: 32768layer 32: GQA: 128 query / 8 KV heads · head 64layer 32: MLP: 32768layer 33: GQA: 128 query / 8 KV heads · head 64layer 33: MLP: 32768layer 34: GQA: 128 query / 8 KV heads · head 64layer 34: MLP: 32768layer 35: GQA: 128 query / 8 KV heads · head 64layer 35: MLP: 32768layer 36: GQA: 128 query / 8 KV heads · head 64layer 36: MLP: 32768layer 37: GQA: 128 query / 8 KV heads · head 64layer 37: MLP: 32768layer 38: GQA: 128 query / 8 KV heads · head 64layer 38: MLP: 32768layer 39: GQA: 128 query / 8 KV heads · head 64layer 39: MLP: 32768layer 40: GQA: 128 query / 8 KV heads · head 64layer 40: MLP: 32768layer 41: GQA: 128 query / 8 KV heads · head 64layer 41: MLP: 32768layer 42: GQA: 128 query / 8 KV heads · head 64layer 42: MLP: 32768layer 43: GQA: 128 query / 8 KV heads · head 64layer 43: MLP: 32768layer 44: GQA: 128 query / 8 KV heads · head 64layer 44: MLP: 32768layer 45: GQA: 128 query / 8 KV heads · head 64layer 45: MLP: 32768layer 46: GQA: 128 query / 8 KV heads · head 64layer 46: MLP: 32768layer 47: GQA: 128 query / 8 KV heads · head 64layer 47: MLP: 32768layer 48: GQA: 128 query / 8 KV heads · head 64layer 48: MLP: 32768layer 49: GQA: 128 query / 8 KV heads · head 64layer 49: MLP: 32768layer 50: GQA: 128 query / 8 KV heads · head 64layer 50: MLP: 32768layer 51: GQA: 128 query / 8 KV heads · head 64layer 51: MLP: 32768layer 52: GQA: 128 query / 8 KV heads · head 64layer 52: MLP: 32768layer 53: GQA: 128 query / 8 KV heads · head 64layer 53: MLP: 32768layer 54: GQA: 128 query / 8 KV heads · head 64layer 54: MLP: 32768layer 55: GQA: 128 query / 8 KV heads · head 64layer 55: MLP: 32768layer 56: GQA: 128 query / 8 KV heads · head 64layer 56: MLP: 32768layer 57: GQA: 128 query / 8 KV heads · head 64layer 57: MLP: 32768layer 58: GQA: 128 query / 8 KV heads · head 64layer 58: MLP: 32768layer 59: GQA: 128 query / 8 KV heads · head 64layer 59: MLP: 3276803059× 60normGQA: 128 query / 8 KV heads · head 64MLP: 32768+full attentiondense FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)41.3B
Active per token (modelled)41.3B
Without embeddings and output head40.8B total, 40.8B active
Published weights (Hugging Face count)41.8B
KV cache per token, BF16 (layers that grow with context)120 KiB
KV cache + state at 128K tokens, BF1615 GiB
Decode FLOPs per token at 4K context90.7 GFLOP
Prefill FLOPs for a 4K prompt350 TFLOP

KV cache against context

Falcon-40B: KV cache bytes against context length101001,00010,000100,000980 KiB9.5 MiB95 MiB950 MiB9.3 GiBcontext (tokens)KV cache + state (BF16)Falcon-40B

Compare with other models →

Every architecture field

FieldValueSource
d_model8,192config.jsonconfig.jsonhidden_size
vocab65,024config.jsonconfig.jsonvocab_size
tied_embeddingstruecodemodelling codemodeling_falcon.py: lm_head tied to word_embeddings
mixers.full.typeattncodemodelling codeattention
mixers.full.heads128config.jsonconfig.jsonnum_attention_heads
mixers.full.kv_heads8config.jsonconfig.jsonnum_kv_heads
mixers.full.head_dim64codemodelling codehidden_size / num_attention_heads
ffns.dense.typedensecodemodelling codeMLP
ffns.dense.d_ff32,768codemodelling codemodeling_falcon.py: 4 * hidden_size
ffns.dense.gatedfalsecodemodelling codeGELU MLP
layout60× full/denseconfig.jsonconfig.jsonnum_hidden_layers, n_layer
norms_per_layer1codemodelling codeparallel attention and MLP

Sources