llm-architectures-explained

/models

Qwen3.6 35B-A3B

Alibaba Cloud (Qwen) · Qwen3.5 · open weights · multimodal (text stack modelled)

Facts and where they come from

Released2026-04config.jsonconfig.jsonHugging Face repository creation date (api.createdAt)
Licenceapache-2.0config.jsonconfig.jsonREADME metadata: license
Total parameters35Blabmodel cardREADME: 35B in total and 3B activated
Active parameters3Blabmodel cardREADME: 3B activated
Context length256K tokenslabmodel cardREADME: 262,144 natively
Norm placementprecodemodelling codetransformers 5.18.0 / repo modelling code for qwen3_5_moe_text: input_layernorm before attention, post_attention_layernorm before the MLP
Norm typeRMSNormcodemodelling codetransformers 5.18.0 / repo modelling code for qwen3_5_moe_text: input_layernorm before attention, post_attention_layernorm before the MLP
QK-normyescodemodelling codetransformers 5.18.0 qwen3_5_moe: q_norm and k_norm
Positional encodingmultimodal RoPE on 25% of each headconfig.jsonconfig.jsontext_config.rope_parameters.partial_rotary_factor
Parallel attention and MLPnocodemodelling codetransformers 5.18.0 / repo modelling code for qwen3_5_moe_text: input_layernorm before attention, post_attention_layernorm before the MLP

Architecture, drawn from the data

30× Gated DeltaNet + 10× GQA 16q/2kv. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

Qwen3.6 35B-A3B: layer stack and blockslayers (40)mixer / FFNlayer 0: Gated DeltaNet: 32 heads · state 128×128 per headlayer 0: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 1: Gated DeltaNet: 32 heads · state 128×128 per headlayer 1: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 2: Gated DeltaNet: 32 heads · state 128×128 per headlayer 2: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 3: GQA: 16 query / 2 KV heads · head 256 · elementwise output gatelayer 3: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 4: Gated DeltaNet: 32 heads · state 128×128 per headlayer 4: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 5: Gated DeltaNet: 32 heads · state 128×128 per headlayer 5: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 6: Gated DeltaNet: 32 heads · state 128×128 per headlayer 6: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 7: GQA: 16 query / 2 KV heads · head 256 · elementwise output gatelayer 7: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 8: Gated DeltaNet: 32 heads · state 128×128 per headlayer 8: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 9: Gated DeltaNet: 32 heads · state 128×128 per headlayer 9: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 10: Gated DeltaNet: 32 heads · state 128×128 per headlayer 10: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 11: GQA: 16 query / 2 KV heads · head 256 · elementwise output gatelayer 11: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 12: Gated DeltaNet: 32 heads · state 128×128 per headlayer 12: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 13: Gated DeltaNet: 32 heads · state 128×128 per headlayer 13: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 14: Gated DeltaNet: 32 heads · state 128×128 per headlayer 14: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 15: GQA: 16 query / 2 KV heads · head 256 · elementwise output gatelayer 15: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 16: Gated DeltaNet: 32 heads · state 128×128 per headlayer 16: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 17: Gated DeltaNet: 32 heads · state 128×128 per headlayer 17: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 18: Gated DeltaNet: 32 heads · state 128×128 per headlayer 18: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 19: GQA: 16 query / 2 KV heads · head 256 · elementwise output gatelayer 19: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 20: Gated DeltaNet: 32 heads · state 128×128 per headlayer 20: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 21: Gated DeltaNet: 32 heads · state 128×128 per headlayer 21: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 22: Gated DeltaNet: 32 heads · state 128×128 per headlayer 22: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 23: GQA: 16 query / 2 KV heads · head 256 · elementwise output gatelayer 23: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 24: Gated DeltaNet: 32 heads · state 128×128 per headlayer 24: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 25: Gated DeltaNet: 32 heads · state 128×128 per headlayer 25: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 26: Gated DeltaNet: 32 heads · state 128×128 per headlayer 26: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 27: GQA: 16 query / 2 KV heads · head 256 · elementwise output gatelayer 27: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 28: Gated DeltaNet: 32 heads · state 128×128 per headlayer 28: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 29: Gated DeltaNet: 32 heads · state 128×128 per headlayer 29: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 30: Gated DeltaNet: 32 heads · state 128×128 per headlayer 30: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 31: GQA: 16 query / 2 KV heads · head 256 · elementwise output gatelayer 31: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 32: Gated DeltaNet: 32 heads · state 128×128 per headlayer 32: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 33: Gated DeltaNet: 32 heads · state 128×128 per headlayer 33: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 34: Gated DeltaNet: 32 heads · state 128×128 per headlayer 34: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 35: GQA: 16 query / 2 KV heads · head 256 · elementwise output gatelayer 35: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 36: Gated DeltaNet: 32 heads · state 128×128 per headlayer 36: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 37: Gated DeltaNet: 32 heads · state 128×128 per headlayer 37: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 38: Gated DeltaNet: 32 heads · state 128×128 per headlayer 38: MoE: 256 experts, 8 active · expert 512 · 1 sharedlayer 39: GQA: 16 query / 2 KV heads · head 256 · elementwise output gatelayer 39: MoE: 256 experts, 8 active · expert 512 · 1 shared02039× 30normGated DeltaNet: 32 heads · state 128×128 per head+normMoE: 256 experts, 8 active · expert 512 · 1 shared+× 10normGQA: 16 query / 2 KV heads · head 256 · elementwise output gate+normMoE: 256 experts, 8 active · expert 512 · 1 shared+Gated DeltaNetfull attentionMoE FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)34.7B
Active per token (modelled)3.45B
Without embeddings and output head33.6B total, 2.44B active
Multi-token-prediction layers (extra)845M
Published weights (Hugging Face count)36B
KV cache per token, BF16 (layers that grow with context)20 KiB
KV cache + state at 256K tokens, BF165.03 GiB
Decode FLOPs per token at 4K context6.66 GFLOP
Prefill FLOPs for a 4K prompt21.7 TFLOP

KV cache against context

Qwen3.6 35B-A3B: KV cache bytes against context length101001,00010,000100,00095 MiB950 MiBcontext (tokens)KV cache + state (BF16)Qwen3.6 35B-A3B

Compare with other models →

Every architecture field

FieldValueSource
d_model2,048config.jsonconfig.jsontext_config.hidden_size
vocab248,320config.jsonconfig.jsontext_config.vocab_size
tied_embeddingsfalseconfig.jsonconfig.jsontext_config.tie_word_embeddings
mixers.full.typeattncodemodelling codeattention block
mixers.full.heads16config.jsonconfig.jsontext_config.num_attention_heads
mixers.full.kv_heads2config.jsonconfig.jsontext_config.num_key_value_heads
mixers.full.head_dim256config.jsonconfig.jsontext_config.head_dim
mixers.full.gateelementwisecodemodelling codeattn_output_gate: q_proj also produces a sigmoid output gate
mixers.full.qk_normtruecodemodelling codetransformers 5.18.0 qwen3_5_moe_text: q_norm and k_norm
mixers.linear.typedeltanetcodemodelling codeGated DeltaNet linear attention
mixers.linear.k_heads16config.jsonconfig.jsontext_config.linear_num_key_heads
mixers.linear.v_heads32config.jsonconfig.jsontext_config.linear_num_value_heads
mixers.linear.k_head_dim128config.jsonconfig.jsontext_config.linear_key_head_dim
mixers.linear.v_head_dim128config.jsonconfig.jsontext_config.linear_value_head_dim
mixers.linear.conv_kernel4config.jsonconfig.jsontext_config.linear_conv_kernel_dim
ffns.moe.typemoecodemodelling codeMoE block
ffns.moe.experts256config.jsonconfig.jsontext_config.num_experts
ffns.moe.active8config.jsonconfig.jsontext_config.num_experts_per_tok
ffns.moe.d_expert512config.jsonconfig.jsontext_config.moe_intermediate_size
ffns.moe.gatedtruecodemodelling codeexperts are gated MLPs
ffns.moe.shared1codemodelling codeone shared expert of shared_expert_intermediate_size
ffns.moe.d_shared512config.jsonconfig.jsontext_config.shared_expert_intermediate_size
layout3× linear/moe · 1× full/moe · 3× linear/moe · 1× full/moe · 3× linear/moe · 1× full/moe · 3× linear/moe · 1× full/moe · 3× linear/moe · 1× full/moe · 3× linear/moe · 1× full/moe · 3× linear/moe · 1× full/moe · 3× linear/moe · 1× full/moe · 3× linear/moe · 1× full/moe · 3× linear/moe · 1× full/moeconfig.jsonconfig.jsontext_config.layer_types
mtp_layers1config.jsonconfig.jsontext_config.mtp_num_hidden_layers
mtp_layer{"mixer":"full","ffn":"moe","n":1}codemodelling codeMTP layer: a full-attention block

Sources

Listed in the LLM Architecture Gallery checklist as “Qwen3.6 (35B-A3B)” (name only; see about).