llm-architectures-explained

/models

Llama 4 Maverick

Meta · Llama 4 · open weights · multimodal (text stack modelled)

Facts and where they come from

Released2025-04config.jsonconfig.jsonHugging Face repository creation date (api.createdAt)
Licenceotherconfig.jsonconfig.jsonREADME metadata: license
Total parameters400Blabmodel cardREADME table: 400B (Total)
Active parameters17Blabmodel cardREADME table: 17B (Activated)
Context length1,000,000 tokenslabmodel cardREADME table: context length 1M
Experts128labmodel cardREADME: Llama 4 Maverick (17Bx128E)
Norm placementprecodemodelling codetransformers 5.18.0 / repo modelling code for llama4_text: input_layernorm before attention, post_attention_layernorm before the MLP
Norm typeRMSNormcodemodelling codetransformers 5.18.0 / repo modelling code for llama4_text: input_layernorm before attention, post_attention_layernorm before the MLP
QK-normnoconfig.jsonconfig.jsontext_config.use_qk_norm
Positional encodingRoPE; every 4th layer has no RoPE (global attention); the others use RoPE within their attention chunkcodemodelling codetransformers 5.18.0 llama4: no_rope_layer_interval default 4; RoPE layers are chunked_attention
Parallel attention and MLPnocodemodelling codetransformers 5.18.0 / repo modelling code for llama4_text: input_layernorm before attention, post_attention_layernorm before the MLP

Architecture, drawn from the data

36× GQA 40q/8kv, chunks of 8192 + 12× GQA 40q/8kv. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

Llama 4 Maverick: layer stack and blockslayers (48)mixer / FFNlayer 0: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 0: Gated MLP: 16384layer 1: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 1: MoE: 128 experts, 1 active · expert 8192 · 1 sharedlayer 2: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 2: Gated MLP: 16384layer 3: GQA: 40 query / 8 KV heads · head 128layer 3: MoE: 128 experts, 1 active · expert 8192 · 1 sharedlayer 4: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 4: Gated MLP: 16384layer 5: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 5: MoE: 128 experts, 1 active · expert 8192 · 1 sharedlayer 6: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 6: Gated MLP: 16384layer 7: GQA: 40 query / 8 KV heads · head 128layer 7: MoE: 128 experts, 1 active · expert 8192 · 1 sharedlayer 8: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 8: Gated MLP: 16384layer 9: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 9: MoE: 128 experts, 1 active · expert 8192 · 1 sharedlayer 10: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 10: Gated MLP: 16384layer 11: GQA: 40 query / 8 KV heads · head 128layer 11: MoE: 128 experts, 1 active · expert 8192 · 1 sharedlayer 12: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 12: Gated MLP: 16384layer 13: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 13: MoE: 128 experts, 1 active · expert 8192 · 1 sharedlayer 14: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 14: Gated MLP: 16384layer 15: GQA: 40 query / 8 KV heads · head 128layer 15: MoE: 128 experts, 1 active · expert 8192 · 1 sharedlayer 16: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 16: Gated MLP: 16384layer 17: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 17: MoE: 128 experts, 1 active · expert 8192 · 1 sharedlayer 18: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 18: Gated MLP: 16384layer 19: GQA: 40 query / 8 KV heads · head 128layer 19: MoE: 128 experts, 1 active · expert 8192 · 1 sharedlayer 20: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 20: Gated MLP: 16384layer 21: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 21: MoE: 128 experts, 1 active · expert 8192 · 1 sharedlayer 22: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 22: Gated MLP: 16384layer 23: GQA: 40 query / 8 KV heads · head 128layer 23: MoE: 128 experts, 1 active · expert 8192 · 1 sharedlayer 24: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 24: Gated MLP: 16384layer 25: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 25: MoE: 128 experts, 1 active · expert 8192 · 1 sharedlayer 26: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 26: Gated MLP: 16384layer 27: GQA: 40 query / 8 KV heads · head 128layer 27: MoE: 128 experts, 1 active · expert 8192 · 1 sharedlayer 28: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 28: Gated MLP: 16384layer 29: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 29: MoE: 128 experts, 1 active · expert 8192 · 1 sharedlayer 30: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 30: Gated MLP: 16384layer 31: GQA: 40 query / 8 KV heads · head 128layer 31: MoE: 128 experts, 1 active · expert 8192 · 1 sharedlayer 32: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 32: Gated MLP: 16384layer 33: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 33: MoE: 128 experts, 1 active · expert 8192 · 1 sharedlayer 34: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 34: Gated MLP: 16384layer 35: GQA: 40 query / 8 KV heads · head 128layer 35: MoE: 128 experts, 1 active · expert 8192 · 1 sharedlayer 36: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 36: Gated MLP: 16384layer 37: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 37: MoE: 128 experts, 1 active · expert 8192 · 1 sharedlayer 38: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 38: Gated MLP: 16384layer 39: GQA: 40 query / 8 KV heads · head 128layer 39: MoE: 128 experts, 1 active · expert 8192 · 1 sharedlayer 40: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 40: Gated MLP: 16384layer 41: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 41: MoE: 128 experts, 1 active · expert 8192 · 1 sharedlayer 42: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 42: Gated MLP: 16384layer 43: GQA: 40 query / 8 KV heads · head 128layer 43: MoE: 128 experts, 1 active · expert 8192 · 1 sharedlayer 44: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 44: Gated MLP: 16384layer 45: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 45: MoE: 128 experts, 1 active · expert 8192 · 1 sharedlayer 46: GQA: 40 query / 8 KV heads · head 128 · chunks of 8,192layer 46: Gated MLP: 16384layer 47: GQA: 40 query / 8 KV heads · head 128layer 47: MoE: 128 experts, 1 active · expert 8192 · 1 shared02447× 24normGQA: 40 query / 8 KV heads · head 128 · chunks of 8,192+normGated MLP: 16384+× 12normGQA: 40 query / 8 KV heads · head 128 · chunks of 8,192+normMoE: 128 experts, 1 active · expert 8192 · 1 shared+× 12normGQA: 40 query / 8 KV heads · head 128+normMoE: 128 experts, 1 active · expert 8192 · 1 shared+chunked attentionfull attentiondense FFNMoE FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)401B
Active per token (modelled)17.2B
Without embeddings and output head399B total, 15.1B active
Published weights (Hugging Face count)402B
KV cache per token, BF16 (layers that grow with context)48 KiB
KV cache + state at 1,000,000 tokens, BF1646.9 GiB
Decode FLOPs per token at 4K context36.3 GFLOP
Prefill FLOPs for a 4K prompt132 TFLOP

KV cache against context

Llama 4 Maverick: KV cache bytes against context length101001,00010,000100,000980 KiB9.5 MiB95 MiB950 MiB9.3 GiBcontext (tokens)KV cache + state (BF16)Llama 4 Maverick

Compare with other models →

Every architecture field

FieldValueSource
d_model5,120config.jsonconfig.jsontext_config.hidden_size
vocab202,048config.jsonconfig.jsontext_config.vocab_size
tied_embeddingsfalsecodemodelling codetransformers 5.18.0: PretrainedConfig.tie_word_embeddings defaultnot in config
mixers.full.typeattncodemodelling codeattention block
mixers.full.heads40config.jsonconfig.jsontext_config.num_attention_heads
mixers.full.kv_heads8config.jsonconfig.jsontext_config.num_key_value_heads
mixers.full.head_dim128config.jsonconfig.jsontext_config.head_dim
mixers.chunked.typeattncodemodelling codeattention block
mixers.chunked.heads40config.jsonconfig.jsontext_config.num_attention_heads
mixers.chunked.kv_heads8config.jsonconfig.jsontext_config.num_key_value_heads
mixers.chunked.head_dim128config.jsonconfig.jsontext_config.head_dim
mixers.chunked.chunk8,192config.jsonconfig.jsontext_config.attention_chunk_size
ffns.dense.typedensecodemodelling codeMLP block
ffns.dense.d_ff16,384config.jsonconfig.jsontext_config.intermediate_size_mlp
ffns.dense.gatedtruecodemodelling codeMLP: gated (SwiGLU/GeGLU)
ffns.moe.typemoecodemodelling codeMoE block
ffns.moe.experts128config.jsonconfig.jsontext_config.num_local_experts
ffns.moe.active1config.jsonconfig.jsontext_config.num_experts_per_tok
ffns.moe.d_expert8,192config.jsonconfig.jsontext_config.intermediate_size
ffns.moe.gatedtruecodemodelling codeexperts are gated MLPs
ffns.moe.shared1codemodelling codetransformers 5.18.0 llama4: Llama4TextMoe.shared_expert, one Llama4TextMLP of intermediate_size
ffns.moe.d_shared8,192config.jsonconfig.jsontext_config.intermediate_size
layout1× chunked/dense · 1× chunked/moe · 1× chunked/dense · 1× full/moe · 1× chunked/dense · 1× chunked/moe · 1× chunked/dense · 1× full/moe · 1× chunked/dense · 1× chunked/moe · 1× chunked/dense · 1× full/moe · 1× chunked/dense · 1× chunked/moe · 1× chunked/dense · 1× full/moe · 1× chunked/dense · 1× chunked/moe · 1× chunked/dense · 1× full/moe · 1× chunked/dense · 1× chunked/moe · 1× chunked/dense · 1× full/moe · 1× chunked/dense · 1× chunked/moe · 1× chunked/dense · 1× full/moe · 1× chunked/dense · 1× chunked/moe · 1× chunked/dense · 1× full/moe · 1× chunked/dense · 1× chunked/moe · 1× chunked/dense · 1× full/moe · 1× chunked/dense · 1× chunked/moe · 1× chunked/dense · 1× full/moe · 1× chunked/dense · 1× chunked/moe · 1× chunked/dense · 1× full/moe · 1× chunked/dense · 1× chunked/moe · 1× chunked/dense · 1× full/moecodemodelling codetransformers 5.18.0 llama4: no_rope_layers default (layer i uses RoPE and chunked_attention unless (i + 1) % no_rope_layer_interval == 0; interval 4); moe_layers default range(interleave_moe_layer_step - 1, num_hidden_layers, interleave_moe_layer_step) with text_config.interleave_moe_layer_step = 2

Sources

Listed in the LLM Architecture Gallery checklist as “Llama 4 Maverick (400B)” (name only; see about).