llm-architectures-explained

/models

Llama 3.2 1B

Meta · Llama 3 · open weights

Facts and where they come from

Released2024-09config.jsonconfig.jsonHugging Face repository creation date (api.createdAt)
Licencellama3.2config.jsonconfig.jsonREADME metadata: license
Total parameters1.23Blabmodel cardREADME table: 1B (1.23B)
Active parametersnot disclosednot disclosed
Context length128K tokenslabmeta-llama/llama-models sku_list.py (pinned)model card (README): context length 128k
Norm placementprecodemodelling codetransformers 5.18.0 / repo modelling code for llama: input_layernorm before attention, post_attention_layernorm before the MLP
Norm typeRMSNormcodemodelling codetransformers 5.18.0 / repo modelling code for llama: input_layernorm before attention, post_attention_layernorm before the MLP
QK-normnocodemodelling codeno q/k normalisation in the attention block
Positional encodingRoPEcodemodelling coderotary on the full head (default)
Parallel attention and MLPnocodemodelling codetransformers 5.18.0 / repo modelling code for llama: input_layernorm before attention, post_attention_layernorm before the MLP

Architecture, drawn from the data

GQA 32q/8kv. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

Llama 3.2 1B: layer stack and blockslayers (16)mixer / FFNlayer 0: GQA: 32 query / 8 KV heads · head 64layer 0: Gated MLP: 8192layer 1: GQA: 32 query / 8 KV heads · head 64layer 1: Gated MLP: 8192layer 2: GQA: 32 query / 8 KV heads · head 64layer 2: Gated MLP: 8192layer 3: GQA: 32 query / 8 KV heads · head 64layer 3: Gated MLP: 8192layer 4: GQA: 32 query / 8 KV heads · head 64layer 4: Gated MLP: 8192layer 5: GQA: 32 query / 8 KV heads · head 64layer 5: Gated MLP: 8192layer 6: GQA: 32 query / 8 KV heads · head 64layer 6: Gated MLP: 8192layer 7: GQA: 32 query / 8 KV heads · head 64layer 7: Gated MLP: 8192layer 8: GQA: 32 query / 8 KV heads · head 64layer 8: Gated MLP: 8192layer 9: GQA: 32 query / 8 KV heads · head 64layer 9: Gated MLP: 8192layer 10: GQA: 32 query / 8 KV heads · head 64layer 10: Gated MLP: 8192layer 11: GQA: 32 query / 8 KV heads · head 64layer 11: Gated MLP: 8192layer 12: GQA: 32 query / 8 KV heads · head 64layer 12: Gated MLP: 8192layer 13: GQA: 32 query / 8 KV heads · head 64layer 13: Gated MLP: 8192layer 14: GQA: 32 query / 8 KV heads · head 64layer 14: Gated MLP: 8192layer 15: GQA: 32 query / 8 KV heads · head 64layer 15: Gated MLP: 81920815× 16normGQA: 32 query / 8 KV heads · head 64+normGated MLP: 8192+full attentiondense FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)1.24B
Active per token (modelled)1.24B
Without embeddings and output head973M total, 973M active
Published weights (Hugging Face count)1.24B
KV cache per token, BF16 (layers that grow with context)32 KiB
KV cache + state at 128K tokens, BF164 GiB
Decode FLOPs per token at 4K context3.01 GFLOP
Prefill FLOPs for a 4K prompt9.07 TFLOP

KV cache against context

Llama 3.2 1B: KV cache bytes against context length101001,00010,000100,00098 KiB980 KiB9.5 MiB95 MiB950 MiBcontext (tokens)KV cache + state (BF16)Llama 3.2 1B

Compare with other models →

Every architecture field

FieldValueSource
d_model2,048labmeta-llama/llama-models sku_list.py (pinned)sku_list.py arch_args_1b: dim
vocab128,256labmeta-llama/llama-models sku_list.py (pinned)sku_list.py: LLAMA3_VOCAB_SIZE
tied_embeddingstruelabmeta-llama/llama-models sku_list.py (pinned)model card: shared embeddings
mixers.full.typeattncodemodelling codeattention block
mixers.full.heads32labmeta-llama/llama-models sku_list.py (pinned)sku_list.py arch_args_1b: n_heads
mixers.full.kv_heads8labmeta-llama/llama-models sku_list.py (pinned)sku_list.py arch_args_1b: n_kv_heads
mixers.full.head_dim64codemodelling codetransformers 5.18.0: head_dim = hidden_size / num_attention_heads
ffns.dense.typedensecodemodelling codeMLP block
ffns.dense.d_ff8,192labmeta-llama/llama-models sku_list.py (pinned)llama3/model.py FeedForward: int(2*4*dim/3), x ffn_dim_multiplier 1.5, rounded up to multiple_of 256
ffns.dense.gatedtruecodemodelling codeMLP: gated (SwiGLU/GeGLU)
layout16× full/denselabmeta-llama/llama-models sku_list.py (pinned)sku_list.py arch_args_1b: n_layers

Sources

Listed in the LLM Architecture Gallery checklist as “Llama 3.2 (1B)” (name only; see about).