llm-architectures-explained

/models

Gemma 4 31B

Google · Gemma 4 · open weights · multimodal (text stack modelled)

Facts and where they come from

Released2026-03config.jsonconfig.jsonHugging Face repository creation date (api.createdAt)
Licenceapache-2.0config.jsonconfig.jsonREADME metadata: license
Total parameters30.7Blabmodel cardREADME: Total Parameters 30.7B
Active parametersnot disclosednot disclosed
Context length256K tokenslabmodel cardREADME: 256K tokens
Norm placementsandwichcodemodelling codetransformers 5.18.0 gemma4: pre and post norms around attention and MLP
Norm typeRMSNormcodemodelling codetransformers 5.18.0 gemma4: pre and post norms around attention and MLP
QK-normyescodemodelling codetransformers 5.18.0 gemma4: q_norm and k_norm
Positional encodingRoPE on 25% of each headconfig.jsonconfig.jsontext_config.rope_parameters.full_attention.partial_rotary_factor (global layers)
Parallel attention and MLPnocodemodelling codetransformers 5.18.0 gemma4: pre and post norms around attention and MLP

Architecture, drawn from the data

50× GQA 32q/16kv, window 1024 + 10× GQA 32q/4kv. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

Gemma 4 31B: layer stack and blockslayers (60)mixer / FFNlayer 0: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 0: Gated MLP: 21504layer 1: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 1: Gated MLP: 21504layer 2: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 2: Gated MLP: 21504layer 3: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 3: Gated MLP: 21504layer 4: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 4: Gated MLP: 21504layer 5: GQA: 32 query / 4 KV heads · head 512 · K = Vlayer 5: Gated MLP: 21504layer 6: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 6: Gated MLP: 21504layer 7: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 7: Gated MLP: 21504layer 8: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 8: Gated MLP: 21504layer 9: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 9: Gated MLP: 21504layer 10: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 10: Gated MLP: 21504layer 11: GQA: 32 query / 4 KV heads · head 512 · K = Vlayer 11: Gated MLP: 21504layer 12: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 12: Gated MLP: 21504layer 13: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 13: Gated MLP: 21504layer 14: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 14: Gated MLP: 21504layer 15: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 15: Gated MLP: 21504layer 16: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 16: Gated MLP: 21504layer 17: GQA: 32 query / 4 KV heads · head 512 · K = Vlayer 17: Gated MLP: 21504layer 18: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 18: Gated MLP: 21504layer 19: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 19: Gated MLP: 21504layer 20: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 20: Gated MLP: 21504layer 21: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 21: Gated MLP: 21504layer 22: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 22: Gated MLP: 21504layer 23: GQA: 32 query / 4 KV heads · head 512 · K = Vlayer 23: Gated MLP: 21504layer 24: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 24: Gated MLP: 21504layer 25: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 25: Gated MLP: 21504layer 26: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 26: Gated MLP: 21504layer 27: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 27: Gated MLP: 21504layer 28: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 28: Gated MLP: 21504layer 29: GQA: 32 query / 4 KV heads · head 512 · K = Vlayer 29: Gated MLP: 21504layer 30: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 30: Gated MLP: 21504layer 31: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 31: Gated MLP: 21504layer 32: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 32: Gated MLP: 21504layer 33: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 33: Gated MLP: 21504layer 34: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 34: Gated MLP: 21504layer 35: GQA: 32 query / 4 KV heads · head 512 · K = Vlayer 35: Gated MLP: 21504layer 36: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 36: Gated MLP: 21504layer 37: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 37: Gated MLP: 21504layer 38: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 38: Gated MLP: 21504layer 39: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 39: Gated MLP: 21504layer 40: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 40: Gated MLP: 21504layer 41: GQA: 32 query / 4 KV heads · head 512 · K = Vlayer 41: Gated MLP: 21504layer 42: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 42: Gated MLP: 21504layer 43: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 43: Gated MLP: 21504layer 44: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 44: Gated MLP: 21504layer 45: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 45: Gated MLP: 21504layer 46: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 46: Gated MLP: 21504layer 47: GQA: 32 query / 4 KV heads · head 512 · K = Vlayer 47: Gated MLP: 21504layer 48: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 48: Gated MLP: 21504layer 49: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 49: Gated MLP: 21504layer 50: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 50: Gated MLP: 21504layer 51: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 51: Gated MLP: 21504layer 52: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 52: Gated MLP: 21504layer 53: GQA: 32 query / 4 KV heads · head 512 · K = Vlayer 53: Gated MLP: 21504layer 54: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 54: Gated MLP: 21504layer 55: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 55: Gated MLP: 21504layer 56: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 56: Gated MLP: 21504layer 57: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 57: Gated MLP: 21504layer 58: GQA: 32 query / 16 KV heads · head 256 · window 1,024layer 58: Gated MLP: 21504layer 59: GQA: 32 query / 4 KV heads · head 512 · K = Vlayer 59: Gated MLP: 2150403059× 50normGQA: 32 query / 16 KV heads · head 256 · window 1,024norm+normGated MLP: 21504norm+× 10normGQA: 32 query / 4 KV heads · head 512 · K = Vnorm+normGated MLP: 21504norm+sliding windowfull attentiondense FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)30.7B
Active per token (modelled)30.7B
Without embeddings and output head29.3B total, 29.3B active
Published weights (Hugging Face count)31.3B
KV cache per token, BF16 (layers that grow with context)40 KiB
KV cache + state at 256K tokens, BF1610.8 GiB
Decode FLOPs per token at 4K context65.8 GFLOP
Prefill FLOPs for a 4K prompt251 TFLOP

KV cache against context

Gemma 4 31B: KV cache bytes against context length101001,00010,000100,0009.5 MiB95 MiB950 MiB9.3 GiBcontext (tokens)KV cache + state (BF16)Gemma 4 31B

Compare with other models →

Every architecture field

FieldValueSource
d_model5,376config.jsonconfig.jsontext_config.hidden_size
vocab262,144config.jsonconfig.jsontext_config.vocab_size
tied_embeddingstruecodemodelling codetransformers 5.18.0 gemma4: tie_word_embeddings default true
mixers.full.typeattncodemodelling codeattention
mixers.full.heads32config.jsonconfig.jsontext_config.num_attention_heads
mixers.full.kv_heads4config.jsonconfig.jsontext_config.num_global_key_value_heads
mixers.full.head_dim512config.jsonconfig.jsontext_config.global_head_dim
mixers.full.qk_normtruecodemodelling codetransformers 5.18.0 gemma4: q_norm and k_norm
mixers.full.k_eq_vtrueconfig.jsonconfig.jsontext_config.attention_k_eq_v
mixers.sliding.typeattncodemodelling codeattention block
mixers.sliding.heads32config.jsonconfig.jsontext_config.num_attention_heads
mixers.sliding.kv_heads16config.jsonconfig.jsontext_config.num_key_value_heads
mixers.sliding.head_dim256config.jsonconfig.jsontext_config.head_dim
mixers.sliding.window1,024config.jsonconfig.jsontext_config.sliding_window
mixers.sliding.qk_normtruecodemodelling codetransformers 5.18.0 gemma4: q_norm and k_norm
ffns.dense.typedensecodemodelling codeMLP block
ffns.dense.d_ff21,504config.jsonconfig.jsontext_config.intermediate_size
ffns.dense.gatedtruecodemodelling codeMLP: gated (SwiGLU/GeGLU)
layout5× sliding/dense · 1× full/dense · 5× sliding/dense · 1× full/dense · 5× sliding/dense · 1× full/dense · 5× sliding/dense · 1× full/dense · 5× sliding/dense · 1× full/dense · 5× sliding/dense · 1× full/dense · 5× sliding/dense · 1× full/dense · 5× sliding/dense · 1× full/dense · 5× sliding/dense · 1× full/dense · 5× sliding/dense · 1× full/denseconfig.jsonconfig.jsontext_config.layer_types, num_kv_shared_layers
norms_per_layer4codemodelling codetransformers 5.18.0 gemma4: sandwich norms (pre and post around attention and MLP)

Sources

Listed in the LLM Architecture Gallery checklist as “Gemma 4 (31B)” (name only; see about).