llm-architectures-explained

/models

Olmo 3 32B

Ai2 · OLMo · open weights

Facts and where they come from

Released2025-11config.jsonconfig.jsonHugging Face repository creation date (api.createdAt)
Licenceapache-2.0config.jsonconfig.jsonREADME metadata: license
Total parameters32Blabmodel cardmodel name Olmo-3-32B
Active parametersnot disclosednot disclosed
Context length64K tokensconfig.jsonconfig.jsonmax_position_embeddings
Norm placementpostcodemodelling codetransformers 5.18.0 olmo3: post_attention_layernorm and post_feedforward_layernorm only
Norm typeRMSNormcodemodelling codetransformers 5.18.0 olmo3: post_attention_layernorm and post_feedforward_layernorm only
QK-normyescodemodelling codetransformers 5.18.0 olmo3: q_norm and k_norm
Positional encodingRoPEcodemodelling coderotary on the full head (default)
Parallel attention and MLPnocodemodelling codetransformers 5.18.0 olmo3: post_attention_layernorm and post_feedforward_layernorm only

Architecture, drawn from the data

48× GQA 40q/8kv, window 4096 + 16× GQA 40q/8kv. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

Olmo 3 32B: layer stack and blockslayers (64)mixer / FFNlayer 0: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 0: Gated MLP: 27648layer 1: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 1: Gated MLP: 27648layer 2: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 2: Gated MLP: 27648layer 3: GQA: 40 query / 8 KV heads · head 128layer 3: Gated MLP: 27648layer 4: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 4: Gated MLP: 27648layer 5: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 5: Gated MLP: 27648layer 6: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 6: Gated MLP: 27648layer 7: GQA: 40 query / 8 KV heads · head 128layer 7: Gated MLP: 27648layer 8: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 8: Gated MLP: 27648layer 9: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 9: Gated MLP: 27648layer 10: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 10: Gated MLP: 27648layer 11: GQA: 40 query / 8 KV heads · head 128layer 11: Gated MLP: 27648layer 12: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 12: Gated MLP: 27648layer 13: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 13: Gated MLP: 27648layer 14: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 14: Gated MLP: 27648layer 15: GQA: 40 query / 8 KV heads · head 128layer 15: Gated MLP: 27648layer 16: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 16: Gated MLP: 27648layer 17: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 17: Gated MLP: 27648layer 18: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 18: Gated MLP: 27648layer 19: GQA: 40 query / 8 KV heads · head 128layer 19: Gated MLP: 27648layer 20: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 20: Gated MLP: 27648layer 21: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 21: Gated MLP: 27648layer 22: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 22: Gated MLP: 27648layer 23: GQA: 40 query / 8 KV heads · head 128layer 23: Gated MLP: 27648layer 24: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 24: Gated MLP: 27648layer 25: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 25: Gated MLP: 27648layer 26: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 26: Gated MLP: 27648layer 27: GQA: 40 query / 8 KV heads · head 128layer 27: Gated MLP: 27648layer 28: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 28: Gated MLP: 27648layer 29: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 29: Gated MLP: 27648layer 30: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 30: Gated MLP: 27648layer 31: GQA: 40 query / 8 KV heads · head 128layer 31: Gated MLP: 27648layer 32: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 32: Gated MLP: 27648layer 33: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 33: Gated MLP: 27648layer 34: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 34: Gated MLP: 27648layer 35: GQA: 40 query / 8 KV heads · head 128layer 35: Gated MLP: 27648layer 36: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 36: Gated MLP: 27648layer 37: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 37: Gated MLP: 27648layer 38: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 38: Gated MLP: 27648layer 39: GQA: 40 query / 8 KV heads · head 128layer 39: Gated MLP: 27648layer 40: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 40: Gated MLP: 27648layer 41: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 41: Gated MLP: 27648layer 42: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 42: Gated MLP: 27648layer 43: GQA: 40 query / 8 KV heads · head 128layer 43: Gated MLP: 27648layer 44: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 44: Gated MLP: 27648layer 45: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 45: Gated MLP: 27648layer 46: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 46: Gated MLP: 27648layer 47: GQA: 40 query / 8 KV heads · head 128layer 47: Gated MLP: 27648layer 48: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 48: Gated MLP: 27648layer 49: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 49: Gated MLP: 27648layer 50: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 50: Gated MLP: 27648layer 51: GQA: 40 query / 8 KV heads · head 128layer 51: Gated MLP: 27648layer 52: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 52: Gated MLP: 27648layer 53: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 53: Gated MLP: 27648layer 54: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 54: Gated MLP: 27648layer 55: GQA: 40 query / 8 KV heads · head 128layer 55: Gated MLP: 27648layer 56: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 56: Gated MLP: 27648layer 57: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 57: Gated MLP: 27648layer 58: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 58: Gated MLP: 27648layer 59: GQA: 40 query / 8 KV heads · head 128layer 59: Gated MLP: 27648layer 60: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 60: Gated MLP: 27648layer 61: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 61: Gated MLP: 27648layer 62: GQA: 40 query / 8 KV heads · head 128 · window 4,096layer 62: Gated MLP: 27648layer 63: GQA: 40 query / 8 KV heads · head 128layer 63: Gated MLP: 2764803263× 48GQA: 40 query / 8 KV heads · head 128 · window 4,096norm+Gated MLP: 27648norm+× 16GQA: 40 query / 8 KV heads · head 128norm+Gated MLP: 27648norm+sliding windowfull attentiondense FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)32.2B
Active per token (modelled)32.2B
Without embeddings and output head31.2B total, 31.2B active
Published weights (Hugging Face count)32.2B
KV cache per token, BF16 (layers that grow with context)64 KiB
KV cache + state at 64K tokens, BF164.75 GiB
Decode FLOPs per token at 4K context68.8 GFLOP
Prefill FLOPs for a 4K prompt267 TFLOP

KV cache against context

Olmo 3 32B: KV cache bytes against context length101001,00010,000980 KiB9.5 MiB95 MiB950 MiBcontext (tokens)KV cache + state (BF16)Olmo 3 32B

Compare with other models →

Every architecture field

FieldValueSource
d_model5,120config.jsonconfig.jsonhidden_size
vocab100,278config.jsonconfig.jsonvocab_size
tied_embeddingsfalseconfig.jsonconfig.jsontie_word_embeddings
mixers.full.typeattncodemodelling codeattention block
mixers.full.heads40config.jsonconfig.jsonnum_attention_heads
mixers.full.kv_heads8config.jsonconfig.jsonnum_key_value_heads
mixers.full.head_dim128codemodelling codetransformers 5.18.0: head_dim = hidden_size / num_attention_heads
mixers.full.qk_normtruecodemodelling codetransformers 5.18.0 olmo3: q_norm and k_norm
mixers.sliding.typeattncodemodelling codeattention block
mixers.sliding.heads40config.jsonconfig.jsonnum_attention_heads
mixers.sliding.kv_heads8config.jsonconfig.jsonnum_key_value_heads
mixers.sliding.head_dim128codemodelling codetransformers 5.18.0: head_dim = hidden_size / num_attention_heads
mixers.sliding.window4,096config.jsonconfig.jsonsliding_window
mixers.sliding.qk_normtruecodemodelling codetransformers 5.18.0 olmo3: q_norm and k_norm
ffns.dense.typedensecodemodelling codeMLP block
ffns.dense.d_ff27,648config.jsonconfig.jsonintermediate_size
ffns.dense.gatedtruecodemodelling codeMLP: gated (SwiGLU/GeGLU)
layout3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/denseconfig.jsonconfig.jsonlayer_types

Sources

Listed in the LLM Architecture Gallery checklist as “OLMo 3 (32B)” (name only; see about).