llm-architectures-explained

/models

Gemma 3 270M

Google · Gemma 3 · open weights

Facts and where they come from

Released2025-08config.jsonconfig.jsonHugging Face repository creation date (api.createdAt)
Licencegemmaconfig.jsonconfig.jsonREADME metadata: license
Total parameters270Mlabmodel cardmodel name gemma-3-270m
Active parametersnot disclosednot disclosed
Context length32K tokenslabgoogle-deepmind/gemma _gemma.py (pinned)model card (README): 32K context for the 270M and 1B sizes
Norm placementsandwichcodemodelling codegoogle-deepmind/gemma _gemma.py: use_post_attn_norm and use_post_ffw_norm
Norm typeRMSNormcodemodelling codegoogle-deepmind/gemma _gemma.py: use_post_attn_norm and use_post_ffw_norm
QK-normyesconfig.jsonconfig.jsonuse_qk_norm
Positional encodingRoPEcodemodelling coderotary on the full head (default)
Parallel attention and MLPnocodemodelling codegoogle-deepmind/gemma _gemma.py: use_post_attn_norm and use_post_ffw_norm

Architecture, drawn from the data

15× MQA 4q/1kv, window 512 + 3× MQA 4q/1kv. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

Gemma 3 270M: layer stack and blockslayers (18)mixer / FFNlayer 0: MQA: 4 query / 1 KV heads · head 256 · window 512layer 0: Gated MLP: 2048layer 1: MQA: 4 query / 1 KV heads · head 256 · window 512layer 1: Gated MLP: 2048layer 2: MQA: 4 query / 1 KV heads · head 256 · window 512layer 2: Gated MLP: 2048layer 3: MQA: 4 query / 1 KV heads · head 256 · window 512layer 3: Gated MLP: 2048layer 4: MQA: 4 query / 1 KV heads · head 256 · window 512layer 4: Gated MLP: 2048layer 5: MQA: 4 query / 1 KV heads · head 256layer 5: Gated MLP: 2048layer 6: MQA: 4 query / 1 KV heads · head 256 · window 512layer 6: Gated MLP: 2048layer 7: MQA: 4 query / 1 KV heads · head 256 · window 512layer 7: Gated MLP: 2048layer 8: MQA: 4 query / 1 KV heads · head 256 · window 512layer 8: Gated MLP: 2048layer 9: MQA: 4 query / 1 KV heads · head 256 · window 512layer 9: Gated MLP: 2048layer 10: MQA: 4 query / 1 KV heads · head 256 · window 512layer 10: Gated MLP: 2048layer 11: MQA: 4 query / 1 KV heads · head 256layer 11: Gated MLP: 2048layer 12: MQA: 4 query / 1 KV heads · head 256 · window 512layer 12: Gated MLP: 2048layer 13: MQA: 4 query / 1 KV heads · head 256 · window 512layer 13: Gated MLP: 2048layer 14: MQA: 4 query / 1 KV heads · head 256 · window 512layer 14: Gated MLP: 2048layer 15: MQA: 4 query / 1 KV heads · head 256 · window 512layer 15: Gated MLP: 2048layer 16: MQA: 4 query / 1 KV heads · head 256 · window 512layer 16: Gated MLP: 2048layer 17: MQA: 4 query / 1 KV heads · head 256layer 17: Gated MLP: 20480917× 15normMQA: 4 query / 1 KV heads · head 256 · window 512norm+normGated MLP: 2048norm+× 3normMQA: 4 query / 1 KV heads · head 256norm+normGated MLP: 2048norm+sliding windowfull attentiondense FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)268M
Active per token (modelled)268M
Without embeddings and output head100M total, 100M active
Published weights (Hugging Face count)268M
KV cache per token, BF16 (layers that grow with context)3 KiB
KV cache + state at 32K tokens, BF16104 MiB
Decode FLOPs per token at 4K context618 MFLOP
Prefill FLOPs for a 4K prompt1.05 TFLOP

KV cache against context

Gemma 3 270M: KV cache bytes against context length101001,00010,00098 KiB980 KiB9.5 MiB95 MiBcontext (tokens)KV cache + state (BF16)Gemma 3 270M

Compare with other models →

Every architecture field

FieldValueSource
d_model640labgoogle-deepmind/gemma _gemma.py (pinned)_gemma.py Gemma3_270M: embed_dim
vocab262,144labgoogle-deepmind/gemma _gemma.py (pinned)_gemma.py Gemma3_270M: num_embed
tied_embeddingstruelabgoogle-deepmind/gemma _gemma.py (pinned)_gemma.py Gemma3_270M: embedder shared with the output layer
mixers.full.typeattncodemodelling codeattention block
mixers.full.heads4labgoogle-deepmind/gemma _gemma.py (pinned)_gemma.py Gemma3_270M: num_heads
mixers.full.kv_heads1labgoogle-deepmind/gemma _gemma.py (pinned)_gemma.py Gemma3_270M: num_kv_heads
mixers.full.head_dim256labgoogle-deepmind/gemma _gemma.py (pinned)_gemma.py Gemma3_270M: head_dim
mixers.full.qk_normtruelabgoogle-deepmind/gemma _gemma.py (pinned)_gemma.py Gemma3_270M: use_qk_norm
mixers.sliding.typeattncodemodelling codeattention block
mixers.sliding.heads4labgoogle-deepmind/gemma _gemma.py (pinned)_gemma.py Gemma3_270M: num_heads
mixers.sliding.kv_heads1labgoogle-deepmind/gemma _gemma.py (pinned)_gemma.py Gemma3_270M: num_kv_heads
mixers.sliding.head_dim256labgoogle-deepmind/gemma _gemma.py (pinned)_gemma.py Gemma3_270M: head_dim
mixers.sliding.window512labgoogle-deepmind/gemma _gemma.py (pinned)_gemma.py Gemma3_270M: sliding_window_size
mixers.sliding.qk_normtruelabgoogle-deepmind/gemma _gemma.py (pinned)_gemma.py Gemma3_270M: use_qk_norm
ffns.dense.typedensecodemodelling codeMLP block
ffns.dense.d_ff2,048labgoogle-deepmind/gemma _gemma.py (pinned)_gemma.py Gemma3_270M: hidden_dim
ffns.dense.gatedtruecodemodelling codeMLP: gated (SwiGLU/GeGLU)
layout5× sliding/dense · 1× full/dense · 5× sliding/dense · 1× full/dense · 5× sliding/dense · 1× full/denselabgoogle-deepmind/gemma _gemma.py (pinned)_gemma.py Gemma3_270M: attention_types (local sliding / global pattern)
norms_per_layer4codemodelling codeGemma 2/3: sandwich norms (pre and post around attention and MLP)

Sources

Listed in the LLM Architecture Gallery checklist as “Gemma 3 (270M)” (name only; see about).