llm-architectures-explained

/models

Gemma 4 E2B

Google · Gemma 4 · open weights · multimodal (text stack modelled)

Facts and where they come from

Released2026-03config.jsonconfig.jsonHugging Face repository creation date (api.createdAt)
Licenceapache-2.0config.jsonconfig.jsonREADME metadata: license
Total parameters5.1Blabmodel cardREADME: 2.3B effective (5.1B with embeddings)
Active parametersnot disclosednot disclosed
Context length128K tokenslabmodel cardREADME: 128K tokens
Norm placementsandwichcodemodelling codetransformers 5.18.0 gemma4: pre and post norms around attention and MLP
Norm typeRMSNormcodemodelling codetransformers 5.18.0 gemma4: pre and post norms around attention and MLP
QK-normyescodemodelling codetransformers 5.18.0 gemma4: q_norm and k_norm
Positional encodingRoPE on 25% of each headconfig.jsonconfig.jsontext_config.rope_parameters.full_attention.partial_rotary_factor (global layers)
Parallel attention and MLPnocodemodelling codetransformers 5.18.0 gemma4: pre and post norms around attention and MLP

Architecture, drawn from the data

28× MQA 8q/1kv, window 512 + 7× MQA 8q/1kv. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

Gemma 4 E2B: layer stack and blockslayers (35)mixer / FFNlayer 0: MQA: 8 query / 1 KV heads · head 256 · window 512layer 0: Gated MLP: 6144layer 1: MQA: 8 query / 1 KV heads · head 256 · window 512layer 1: Gated MLP: 6144layer 2: MQA: 8 query / 1 KV heads · head 256 · window 512layer 2: Gated MLP: 6144layer 3: MQA: 8 query / 1 KV heads · head 256 · window 512layer 3: Gated MLP: 6144layer 4: MQA: 8 query / 1 KV heads · head 512layer 4: Gated MLP: 6144layer 5: MQA: 8 query / 1 KV heads · head 256 · window 512layer 5: Gated MLP: 6144layer 6: MQA: 8 query / 1 KV heads · head 256 · window 512layer 6: Gated MLP: 6144layer 7: MQA: 8 query / 1 KV heads · head 256 · window 512layer 7: Gated MLP: 6144layer 8: MQA: 8 query / 1 KV heads · head 256 · window 512layer 8: Gated MLP: 6144layer 9: MQA: 8 query / 1 KV heads · head 512layer 9: Gated MLP: 6144layer 10: MQA: 8 query / 1 KV heads · head 256 · window 512layer 10: Gated MLP: 6144layer 11: MQA: 8 query / 1 KV heads · head 256 · window 512layer 11: Gated MLP: 6144layer 12: MQA: 8 query / 1 KV heads · head 256 · window 512layer 12: Gated MLP: 6144layer 13: MQA: 8 query / 1 KV heads · head 256 · window 512layer 13: Gated MLP: 6144layer 14: MQA: 8 query / 1 KV heads · head 512layer 14: Gated MLP: 6144layer 15: MQA: 8 query / 1 KV heads · head 256 · window 512 (reuses another layer's KV)layer 15: Gated MLP: 12288layer 16: MQA: 8 query / 1 KV heads · head 256 · window 512 (reuses another layer's KV)layer 16: Gated MLP: 12288layer 17: MQA: 8 query / 1 KV heads · head 256 · window 512 (reuses another layer's KV)layer 17: Gated MLP: 12288layer 18: MQA: 8 query / 1 KV heads · head 256 · window 512 (reuses another layer's KV)layer 18: Gated MLP: 12288layer 19: MQA: 8 query / 1 KV heads · head 512 (reuses another layer's KV)layer 19: Gated MLP: 12288layer 20: MQA: 8 query / 1 KV heads · head 256 · window 512 (reuses another layer's KV)layer 20: Gated MLP: 12288layer 21: MQA: 8 query / 1 KV heads · head 256 · window 512 (reuses another layer's KV)layer 21: Gated MLP: 12288layer 22: MQA: 8 query / 1 KV heads · head 256 · window 512 (reuses another layer's KV)layer 22: Gated MLP: 12288layer 23: MQA: 8 query / 1 KV heads · head 256 · window 512 (reuses another layer's KV)layer 23: Gated MLP: 12288layer 24: MQA: 8 query / 1 KV heads · head 512 (reuses another layer's KV)layer 24: Gated MLP: 12288layer 25: MQA: 8 query / 1 KV heads · head 256 · window 512 (reuses another layer's KV)layer 25: Gated MLP: 12288layer 26: MQA: 8 query / 1 KV heads · head 256 · window 512 (reuses another layer's KV)layer 26: Gated MLP: 12288layer 27: MQA: 8 query / 1 KV heads · head 256 · window 512 (reuses another layer's KV)layer 27: Gated MLP: 12288layer 28: MQA: 8 query / 1 KV heads · head 256 · window 512 (reuses another layer's KV)layer 28: Gated MLP: 12288layer 29: MQA: 8 query / 1 KV heads · head 512 (reuses another layer's KV)layer 29: Gated MLP: 12288layer 30: MQA: 8 query / 1 KV heads · head 256 · window 512 (reuses another layer's KV)layer 30: Gated MLP: 12288layer 31: MQA: 8 query / 1 KV heads · head 256 · window 512 (reuses another layer's KV)layer 31: Gated MLP: 12288layer 32: MQA: 8 query / 1 KV heads · head 256 · window 512 (reuses another layer's KV)layer 32: Gated MLP: 12288layer 33: MQA: 8 query / 1 KV heads · head 256 · window 512 (reuses another layer's KV)layer 33: Gated MLP: 12288layer 34: MQA: 8 query / 1 KV heads · head 512 (reuses another layer's KV)layer 34: Gated MLP: 1228801734× 16normMQA: 8 query / 1 KV heads · head 256 · window 512norm+normGated MLP: 12288norm+× 12normMQA: 8 query / 1 KV heads · head 256 · window 512norm+normGated MLP: 6144norm+× 4normMQA: 8 query / 1 KV heads · head 512norm+normGated MLP: 12288norm+× 3normMQA: 8 query / 1 KV heads · head 512norm+normGated MLP: 6144norm+sliding windowfull attentiondense FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)4.59B
Active per token (modelled)4.59B
Without embeddings and output head1.84B total, 1.84B active
Published weights (Hugging Face count)5.12B
KV cache per token, BF16 (layers that grow with context)6 KiB
KV cache + state at 128K tokens, BF16774 MiB
Decode FLOPs per token at 4K context5.06 GFLOP
Prefill FLOPs for a 4K prompt16.5 TFLOP

KV cache against context

Gemma 4 E2B: KV cache bytes against context length101001,00010,000100,00098 KiB980 KiB9.5 MiB95 MiBcontext (tokens)KV cache + state (BF16)Gemma 4 E2B

Compare with other models →

Every architecture field

FieldValueSource
d_model1,536config.jsonconfig.jsontext_config.hidden_size
vocab262,144config.jsonconfig.jsontext_config.vocab_size
tied_embeddingstruecodemodelling codetransformers 5.18.0 gemma4: tie_word_embeddings default true
mixers.full.typeattncodemodelling codeattention
mixers.full.heads8config.jsonconfig.jsontext_config.num_attention_heads
mixers.full.kv_heads1config.jsonconfig.jsontext_config.num_key_value_heads
mixers.full.head_dim512config.jsonconfig.jsontext_config.global_head_dim
mixers.full.qk_normtruecodemodelling codetransformers 5.18.0 gemma4: q_norm and k_norm
mixers.sliding.typeattncodemodelling codeattention block
mixers.sliding.heads8config.jsonconfig.jsontext_config.num_attention_heads
mixers.sliding.kv_heads1config.jsonconfig.jsontext_config.num_key_value_heads
mixers.sliding.head_dim256config.jsonconfig.jsontext_config.head_dim
mixers.sliding.window512config.jsonconfig.jsontext_config.sliding_window
mixers.sliding.qk_normtruecodemodelling codetransformers 5.18.0 gemma4: q_norm and k_norm
ffns.dense.typedensecodemodelling codeMLP block
ffns.dense.d_ff6,144config.jsonconfig.jsontext_config.intermediate_size
ffns.dense.gatedtruecodemodelling codeMLP: gated (SwiGLU/GeGLU)
ffns.dense_wide.typedensecodemodelling codeMLP
ffns.dense_wide.d_ff12,288codemodelling codetransformers 5.18.0 gemma4: use_double_wide_mlp doubles intermediate_size on KV-shared layers
ffns.dense_wide.gatedtruecodemodelling codeGeGLU
layout4× sliding/dense · 1× full/dense · 4× sliding/dense · 1× full/dense · 4× sliding/dense · 1× full/dense · 4× sliding/dense_wide (kv_shared=true) · 1× full/dense_wide (kv_shared=true) · 4× sliding/dense_wide (kv_shared=true) · 1× full/dense_wide (kv_shared=true) · 4× sliding/dense_wide (kv_shared=true) · 1× full/dense_wide (kv_shared=true) · 4× sliding/dense_wide (kv_shared=true) · 1× full/dense_wide (kv_shared=true)config.jsonconfig.jsontext_config.layer_types, num_kv_shared_layers
norms_per_layer4codemodelling codetransformers 5.18.0 gemma4: sandwich norms (pre and post around attention and MLP)
extra_embedding_params2,348,810,240codemodelling codetransformers 5.18.0 gemma4: per-layer embeddings, vocab_size_per_layer_input x num_hidden_layers x hidden_size_per_layer_input

Sources

Listed in the LLM Architecture Gallery checklist as “Gemma 4 (E2B)” (name only; see about).