llm-architectures-explained

/models

Muse Glimmer 30B

Meta · Muse · open weights · multimodal (text stack modelled)

Facts and where they come from

Released2026-08config.jsonconfig.jsonHugging Face repository creation date (api.createdAt)
Licenceapache-2.0config.jsonconfig.jsonREADME metadata: license
Total parameters29.6Blabmodel cardREADME: 29.6B (including vision encoder)
Active parametersnot disclosednot disclosed
Context length128K tokenslabmodel cardREADME: Context length 131,072+
Norm placementsandwichcodemodelling codetransformers 5.18.0 muse_glimmer: pre_feedforward and post_feedforward norms
Norm typeRMSNormcodemodelling codetransformers 5.18.0 muse_glimmer: pre_feedforward and post_feedforward norms
QK-normnocodemodelling codeno q/k normalisation in the attention block
Positional encodingRoPE; global (full-attention) layers have no RoPEconfig.jsonconfig.jsontext_config.layer_rope_theta (0 on global layers)
Parallel attention and MLPnocodemodelling codetransformers 5.18.0 muse_glimmer: pre_feedforward and post_feedforward norms

Architecture, drawn from the data

39× GQA 32q/2kv, window 2048 + 13× GQA 32q/2kv. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

Muse Glimmer 30B: layer stack and blockslayers (52)mixer / FFNlayer 0: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 0: Gated MLP: 19968layer 1: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 1: Gated MLP: 19968layer 2: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 2: Gated MLP: 19968layer 3: GQA: 32 query / 2 KV heads · head 128layer 3: Gated MLP: 19968layer 4: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 4: Gated MLP: 19968layer 5: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 5: Gated MLP: 19968layer 6: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 6: Gated MLP: 19968layer 7: GQA: 32 query / 2 KV heads · head 128layer 7: Gated MLP: 19968layer 8: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 8: Gated MLP: 19968layer 9: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 9: Gated MLP: 19968layer 10: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 10: Gated MLP: 19968layer 11: GQA: 32 query / 2 KV heads · head 128layer 11: Gated MLP: 19968layer 12: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 12: Gated MLP: 19968layer 13: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 13: Gated MLP: 19968layer 14: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 14: Gated MLP: 19968layer 15: GQA: 32 query / 2 KV heads · head 128layer 15: Gated MLP: 19968layer 16: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 16: Gated MLP: 19968layer 17: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 17: Gated MLP: 19968layer 18: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 18: Gated MLP: 19968layer 19: GQA: 32 query / 2 KV heads · head 128layer 19: Gated MLP: 19968layer 20: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 20: Gated MLP: 19968layer 21: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 21: Gated MLP: 19968layer 22: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 22: Gated MLP: 19968layer 23: GQA: 32 query / 2 KV heads · head 128layer 23: Gated MLP: 19968layer 24: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 24: Gated MLP: 19968layer 25: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 25: Gated MLP: 19968layer 26: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 26: Gated MLP: 19968layer 27: GQA: 32 query / 2 KV heads · head 128layer 27: Gated MLP: 19968layer 28: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 28: Gated MLP: 19968layer 29: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 29: Gated MLP: 19968layer 30: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 30: Gated MLP: 19968layer 31: GQA: 32 query / 2 KV heads · head 128layer 31: Gated MLP: 19968layer 32: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 32: Gated MLP: 19968layer 33: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 33: Gated MLP: 19968layer 34: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 34: Gated MLP: 19968layer 35: GQA: 32 query / 2 KV heads · head 128layer 35: Gated MLP: 19968layer 36: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 36: Gated MLP: 19968layer 37: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 37: Gated MLP: 19968layer 38: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 38: Gated MLP: 19968layer 39: GQA: 32 query / 2 KV heads · head 128layer 39: Gated MLP: 19968layer 40: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 40: Gated MLP: 19968layer 41: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 41: Gated MLP: 19968layer 42: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 42: Gated MLP: 19968layer 43: GQA: 32 query / 2 KV heads · head 128layer 43: Gated MLP: 19968layer 44: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 44: Gated MLP: 19968layer 45: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 45: Gated MLP: 19968layer 46: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 46: Gated MLP: 19968layer 47: GQA: 32 query / 2 KV heads · head 128layer 47: Gated MLP: 19968layer 48: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 48: Gated MLP: 19968layer 49: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 49: Gated MLP: 19968layer 50: GQA: 32 query / 2 KV heads · head 128 · window 2,048layer 50: Gated MLP: 19968layer 51: GQA: 32 query / 2 KV heads · head 128layer 51: Gated MLP: 1996802651× 39normGQA: 32 query / 2 KV heads · head 128 · window 2,048norm+normGated MLP: 19968norm+× 13normGQA: 32 query / 2 KV heads · head 128norm+normGated MLP: 19968norm+sliding windowfull attentiondense FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)26.4B
Active per token (modelled)26.4B
Without embeddings and output head23.7B total, 23.7B active
Published weights (Hugging Face count)29.8B
KV cache per token, BF16 (layers that grow with context)13 KiB
KV cache + state at 128K tokens, BF161.7 GiB
Decode FLOPs per token at 4K context52.4 GFLOP
Prefill FLOPs for a 4K prompt200 TFLOP

KV cache against context

Muse Glimmer 30B: KV cache bytes against context length101001,00010,000100,00098 KiB980 KiB9.5 MiB95 MiB950 MiBcontext (tokens)KV cache + state (BF16)Muse Glimmer 30B

Compare with other models →

Every architecture field

FieldValueSource
d_model6,656config.jsonconfig.jsontext_config.hidden_size
vocab202,048config.jsonconfig.jsontext_config.vocab_size
tied_embeddingsfalseconfig.jsonconfig.jsontext_config.tie_word_embeddings
mixers.full.typeattncodemodelling codeattention block
mixers.full.heads32config.jsonconfig.jsontext_config.num_attention_heads
mixers.full.kv_heads2config.jsonconfig.jsontext_config.num_key_value_heads
mixers.full.head_dim128config.jsonconfig.jsontext_config.head_dim
mixers.sliding.typeattncodemodelling codeattention block
mixers.sliding.heads32config.jsonconfig.jsontext_config.num_attention_heads
mixers.sliding.kv_heads2config.jsonconfig.jsontext_config.num_key_value_heads
mixers.sliding.head_dim128config.jsonconfig.jsontext_config.head_dim
mixers.sliding.window2,048config.jsonconfig.jsontext_config.sliding_window
ffns.dense.typedensecodemodelling codeMLP block
ffns.dense.d_ff19,968config.jsonconfig.jsontext_config.intermediate_size
ffns.dense.gatedtruecodemodelling codeMLP: gated (SwiGLU/GeGLU)
layout3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/denseconfig.jsonconfig.jsontext_config.layer_types

Sources

Listed in the LLM Architecture Gallery checklist as “Muse Glimmer (30B)” (name only; see about).