llm-architectures-explained

/models

Command A+

Cohere · Command A · open weights · multimodal (text stack modelled)

Facts and where they come from

Released2026-05config.jsonconfig.jsonHugging Face repository creation date (api.createdAt)
Licenceapache-2.0config.jsonconfig.jsonREADME metadata: license
Total parameters218Blabmodel cardREADME: 25B active parameters, 218B total parameters
Active parameters25Blabmodel cardREADME: 25B active
Context length128K tokenslabmodel cardREADME: context length 128K input
Norm placementparallelcodemodelling codetransformers 5.18.0 cohere2_moe: a single input_layernorm feeds attention and the MoE
Norm typeLayerNormcodemodelling codetransformers 5.18.0 cohere2_moe: a single input_layernorm feeds attention and the MoE
QK-normnoconfig.jsonconfig.jsontext_config.use_qk_norm
Positional encodingRoPE; full-attention layers have no RoPE; sliding-window layers use itcodemodelling codetransformers 5.18.0 cohere2_moe: RoPE applied only when the layer has a sliding window
Parallel attention and MLPyescodemodelling codetransformers 5.18.0 cohere2_moe: a single input_layernorm feeds attention and the MoE

Architecture, drawn from the data

24× GQA 128q/8kv, window 4096 + 8× GQA 128q/8kv. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

Command A+: layer stack and blockslayers (32)mixer / FFNlayer 0: GQA: 128 query / 8 KV heads · head 128 · window 4,096layer 0: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 1: GQA: 128 query / 8 KV heads · head 128 · window 4,096layer 1: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 2: GQA: 128 query / 8 KV heads · head 128 · window 4,096layer 2: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 3: GQA: 128 query / 8 KV heads · head 128layer 3: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 4: GQA: 128 query / 8 KV heads · head 128 · window 4,096layer 4: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 5: GQA: 128 query / 8 KV heads · head 128 · window 4,096layer 5: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 6: GQA: 128 query / 8 KV heads · head 128 · window 4,096layer 6: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 7: GQA: 128 query / 8 KV heads · head 128layer 7: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 8: GQA: 128 query / 8 KV heads · head 128 · window 4,096layer 8: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 9: GQA: 128 query / 8 KV heads · head 128 · window 4,096layer 9: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 10: GQA: 128 query / 8 KV heads · head 128 · window 4,096layer 10: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 11: GQA: 128 query / 8 KV heads · head 128layer 11: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 12: GQA: 128 query / 8 KV heads · head 128 · window 4,096layer 12: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 13: GQA: 128 query / 8 KV heads · head 128 · window 4,096layer 13: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 14: GQA: 128 query / 8 KV heads · head 128 · window 4,096layer 14: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 15: GQA: 128 query / 8 KV heads · head 128layer 15: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 16: GQA: 128 query / 8 KV heads · head 128 · window 4,096layer 16: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 17: GQA: 128 query / 8 KV heads · head 128 · window 4,096layer 17: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 18: GQA: 128 query / 8 KV heads · head 128 · window 4,096layer 18: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 19: GQA: 128 query / 8 KV heads · head 128layer 19: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 20: GQA: 128 query / 8 KV heads · head 128 · window 4,096layer 20: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 21: GQA: 128 query / 8 KV heads · head 128 · window 4,096layer 21: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 22: GQA: 128 query / 8 KV heads · head 128 · window 4,096layer 22: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 23: GQA: 128 query / 8 KV heads · head 128layer 23: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 24: GQA: 128 query / 8 KV heads · head 128 · window 4,096layer 24: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 25: GQA: 128 query / 8 KV heads · head 128 · window 4,096layer 25: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 26: GQA: 128 query / 8 KV heads · head 128 · window 4,096layer 26: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 27: GQA: 128 query / 8 KV heads · head 128layer 27: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 28: GQA: 128 query / 8 KV heads · head 128 · window 4,096layer 28: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 29: GQA: 128 query / 8 KV heads · head 128 · window 4,096layer 29: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 30: GQA: 128 query / 8 KV heads · head 128 · window 4,096layer 30: MoE: 128 experts, 8 active · expert 4096 · 4 sharedlayer 31: GQA: 128 query / 8 KV heads · head 128layer 31: MoE: 128 experts, 8 active · expert 4096 · 4 shared01631× 24normGQA: 128 query / 8 KV heads · head 128 · window 4,096MoE: 128 experts, 8 active · expert 4096 · 4 shared+× 8normGQA: 128 query / 8 KV heads · head 128MoE: 128 experts, 8 active · expert 4096 · 4 shared+sliding windowfull attentionMoE FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)218B
Active per token (modelled)25B
Without embeddings and output head217B total, 23.9B active
Published weights (Hugging Face count)219B (packed low-bit tensors, so not comparable)
KV cache per token, BF16 (layers that grow with context)32 KiB
KV cache + state at 128K tokens, BF164.38 GiB
Decode FLOPs per token at 4K context58.6 GFLOP
Prefill FLOPs for a 4K prompt213 TFLOP

KV cache against context

Command A+: KV cache bytes against context length101001,00010,000100,000980 KiB9.5 MiB95 MiB950 MiBcontext (tokens)KV cache + state (BF16)Command A+

Compare with other models →

Every architecture field

FieldValueSource
d_model4,096config.jsonconfig.jsontext_config.hidden_size
vocab262,144config.jsonconfig.jsontext_config.vocab_size
tied_embeddingstrueconfig.jsonconfig.jsontext_config.use_embedding_sharing
mixers.full.typeattncodemodelling codeattention block
mixers.full.heads128config.jsonconfig.jsontext_config.num_attention_heads
mixers.full.kv_heads8config.jsonconfig.jsontext_config.num_key_value_heads
mixers.full.head_dim128config.jsonconfig.jsontext_config.head_dim
mixers.sliding.typeattncodemodelling codeattention block
mixers.sliding.heads128config.jsonconfig.jsontext_config.num_attention_heads
mixers.sliding.kv_heads8config.jsonconfig.jsontext_config.num_key_value_heads
mixers.sliding.head_dim128config.jsonconfig.jsontext_config.head_dim
mixers.sliding.window4,096config.jsonconfig.jsontext_config.sliding_window
ffns.dense.typedensecodemodelling codeMLP block
ffns.dense.d_ff16,384config.jsonconfig.jsontext_config.prefix_dense_intermediate_size
ffns.dense.gatedtruecodemodelling codeMLP: gated (SwiGLU/GeGLU)
ffns.moe.typemoecodemodelling codeMoE block
ffns.moe.experts128config.jsonconfig.jsontext_config.num_experts
ffns.moe.active8config.jsonconfig.jsontext_config.num_experts_per_tok
ffns.moe.d_expert4,096config.jsonconfig.jsontext_config.intermediate_size
ffns.moe.gatedtruecodemodelling codeexperts are gated MLPs
ffns.moe.shared4config.jsonconfig.jsontext_config.num_shared_experts
ffns.moe.d_shared4,096config.jsonconfig.jsontext_config.intermediate_size
layout3× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moeconfig.jsonconfig.jsontext_config.layer_types, first_k_dense_replace
norms_per_layer1codemodelling codetransformers 5.18.0 cohere2_moe: one LayerNorm per parallel block

Sources

Listed in the LLM Architecture Gallery checklist as “Command A+ (218B-A25B)” (name only; see about).