llm-architectures-explained

/models

Nemotron 3.5 Lightning

NVIDIA · Nemotron 3 · open weights

Facts and where they come from

Released2026-08config.jsonconfig.jsonHugging Face repository creation date (api.createdAt)
Licenceotherconfig.jsonconfig.jsonREADME metadata: license
Total parameters30Blabmodel cardREADME: 30B Total / 3B Active
Active parameters3Blabmodel cardREADME: 3B Active
Context length1M tokenslabmodel cardREADME: up to 1M tokens
Norm placementprecodemodelling codetransformers 5.18.0 nemotron_h: one norm before each mixer or MLP block
Norm typeRMSNormcodemodelling codetransformers 5.18.0 nemotron_h: one norm before each mixer or MLP block
QK-normnocodemodelling codeno q/k normalisation in the attention block
Positional encodingRoPEconfig.jsonconfig.jsonpartial_rotary_factor
Parallel attention and MLPnocodemodelling codetransformers 5.18.0 nemotron_h: one norm before each mixer or MLP block

Architecture, drawn from the data

23× Mamba-2 + 6× GQA 32q/2kv. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

Nemotron 3.5 Lightning: layer stack and blockslayers (52)mixer / FFNlayer 0: Mamba-2: 64 heads × 64 · state 128layer 0: no FFNlayer 1: no mixerlayer 1: MoE: 128 experts, 6 active · expert 1856 · 1 sharedlayer 2: Mamba-2: 64 heads × 64 · state 128layer 2: no FFNlayer 3: no mixerlayer 3: MoE: 128 experts, 6 active · expert 1856 · 1 sharedlayer 4: Mamba-2: 64 heads × 64 · state 128layer 4: no FFNlayer 5: GQA: 32 query / 2 KV heads · head 128layer 5: no FFNlayer 6: no mixerlayer 6: MoE: 128 experts, 6 active · expert 1856 · 1 sharedlayer 7: Mamba-2: 64 heads × 64 · state 128layer 7: no FFNlayer 8: no mixerlayer 8: MoE: 128 experts, 6 active · expert 1856 · 1 sharedlayer 9: Mamba-2: 64 heads × 64 · state 128layer 9: no FFNlayer 10: no mixerlayer 10: MoE: 128 experts, 6 active · expert 1856 · 1 sharedlayer 11: Mamba-2: 64 heads × 64 · state 128layer 11: no FFNlayer 12: GQA: 32 query / 2 KV heads · head 128layer 12: no FFNlayer 13: no mixerlayer 13: MoE: 128 experts, 6 active · expert 1856 · 1 sharedlayer 14: Mamba-2: 64 heads × 64 · state 128layer 14: no FFNlayer 15: no mixerlayer 15: MoE: 128 experts, 6 active · expert 1856 · 1 sharedlayer 16: Mamba-2: 64 heads × 64 · state 128layer 16: no FFNlayer 17: no mixerlayer 17: MoE: 128 experts, 6 active · expert 1856 · 1 sharedlayer 18: Mamba-2: 64 heads × 64 · state 128layer 18: no FFNlayer 19: GQA: 32 query / 2 KV heads · head 128layer 19: no FFNlayer 20: no mixerlayer 20: MoE: 128 experts, 6 active · expert 1856 · 1 sharedlayer 21: Mamba-2: 64 heads × 64 · state 128layer 21: no FFNlayer 22: no mixerlayer 22: MoE: 128 experts, 6 active · expert 1856 · 1 sharedlayer 23: Mamba-2: 64 heads × 64 · state 128layer 23: no FFNlayer 24: no mixerlayer 24: MoE: 128 experts, 6 active · expert 1856 · 1 sharedlayer 25: Mamba-2: 64 heads × 64 · state 128layer 25: no FFNlayer 26: GQA: 32 query / 2 KV heads · head 128layer 26: no FFNlayer 27: no mixerlayer 27: MoE: 128 experts, 6 active · expert 1856 · 1 sharedlayer 28: Mamba-2: 64 heads × 64 · state 128layer 28: no FFNlayer 29: no mixerlayer 29: MoE: 128 experts, 6 active · expert 1856 · 1 sharedlayer 30: Mamba-2: 64 heads × 64 · state 128layer 30: no FFNlayer 31: no mixerlayer 31: MoE: 128 experts, 6 active · expert 1856 · 1 sharedlayer 32: Mamba-2: 64 heads × 64 · state 128layer 32: no FFNlayer 33: GQA: 32 query / 2 KV heads · head 128layer 33: no FFNlayer 34: no mixerlayer 34: MoE: 128 experts, 6 active · expert 1856 · 1 sharedlayer 35: Mamba-2: 64 heads × 64 · state 128layer 35: no FFNlayer 36: no mixerlayer 36: MoE: 128 experts, 6 active · expert 1856 · 1 sharedlayer 37: Mamba-2: 64 heads × 64 · state 128layer 37: no FFNlayer 38: no mixerlayer 38: MoE: 128 experts, 6 active · expert 1856 · 1 sharedlayer 39: Mamba-2: 64 heads × 64 · state 128layer 39: no FFNlayer 40: no mixerlayer 40: MoE: 128 experts, 6 active · expert 1856 · 1 sharedlayer 41: Mamba-2: 64 heads × 64 · state 128layer 41: no FFNlayer 42: GQA: 32 query / 2 KV heads · head 128layer 42: no FFNlayer 43: no mixerlayer 43: MoE: 128 experts, 6 active · expert 1856 · 1 sharedlayer 44: Mamba-2: 64 heads × 64 · state 128layer 44: no FFNlayer 45: no mixerlayer 45: MoE: 128 experts, 6 active · expert 1856 · 1 sharedlayer 46: Mamba-2: 64 heads × 64 · state 128layer 46: no FFNlayer 47: no mixerlayer 47: MoE: 128 experts, 6 active · expert 1856 · 1 sharedlayer 48: Mamba-2: 64 heads × 64 · state 128layer 48: no FFNlayer 49: no mixerlayer 49: MoE: 128 experts, 6 active · expert 1856 · 1 sharedlayer 50: Mamba-2: 64 heads × 64 · state 128layer 50: no FFNlayer 51: no mixerlayer 51: MoE: 128 experts, 6 active · expert 1856 · 1 shared02651× 23normMamba-2: 64 heads × 64 · state 128+× 23normMoE: 128 experts, 6 active · expert 1856 · 1 shared+× 6normGQA: 32 query / 2 KV heads · head 128+Mamba-2full attentionMoE FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)31.6B
Active per token (modelled)3.58B
Without embeddings and output head30.9B total, 2.88B active
Multi-token-prediction layers (extra)1.34B
Published weights (Hugging Face count)17.8B (packed low-bit tensors, so not comparable)
KV cache per token, BF16 (layers that grow with context)6 KiB
KV cache + state at 1M tokens, BF166.02 GiB
Decode FLOPs per token at 4K context6.93 GFLOP
Prefill FLOPs for a 4K prompt24.7 TFLOP

KV cache against context

Nemotron 3.5 Lightning: KV cache bytes against context length101001,00010,000100,0001,000,00095 MiB950 MiBcontext (tokens)KV cache + state (BF16)Nemotron 3.5 Lightning

Compare with other models →

Every architecture field

FieldValueSource
d_model2,688config.jsonconfig.jsonhidden_size
vocab131,072config.jsonconfig.jsonvocab_size
tied_embeddingsfalseconfig.jsonconfig.jsontie_word_embeddings
mixers.full.typeattncodemodelling codeattention block
mixers.full.heads32config.jsonconfig.jsonnum_attention_heads
mixers.full.kv_heads2config.jsonconfig.jsonnum_key_value_heads
mixers.full.head_dim128config.jsonconfig.jsonhead_dim
mixers.mamba.typemamba2codemodelling codeMamba-2 mixer
mixers.mamba.heads64config.jsonconfig.jsonmamba_num_heads
mixers.mamba.head_dim64config.jsonconfig.jsonmamba_head_dim
mixers.mamba.state128config.jsonconfig.jsonssm_state_size
mixers.mamba.groups8config.jsonconfig.jsonn_groups
mixers.mamba.conv_kernel4config.jsonconfig.jsonconv_kernel
ffns.moe.typemoecodemodelling codeMoE block
ffns.moe.experts128config.jsonconfig.jsonn_routed_experts
ffns.moe.active6config.jsonconfig.jsonnum_experts_per_tok
ffns.moe.d_expert1,856config.jsonconfig.jsonmoe_intermediate_size
ffns.moe.gatedfalsecodemodelling codeexperts are two-matrix MLPs
ffns.moe.shared1config.jsonconfig.jsonn_shared_experts
ffns.moe.d_shared3,712config.jsonconfig.jsonmoe_shared_expert_intermediate_size
ffns.dense.typedensecodemodelling codeMLP block
ffns.dense.d_ff1,856config.jsonconfig.jsonintermediate_size
ffns.dense.gatedfalsecodemodelling codeMLP: two matrices, no gate
layout1× mamba/none · 1× none/moe · 1× mamba/none · 1× none/moe · 1× mamba/none · 1× full/none · 1× none/moe · 1× mamba/none · 1× none/moe · 1× mamba/none · 1× none/moe · 1× mamba/none · 1× full/none · 1× none/moe · 1× mamba/none · 1× none/moe · 1× mamba/none · 1× none/moe · 1× mamba/none · 1× full/none · 1× none/moe · 1× mamba/none · 1× none/moe · 1× mamba/none · 1× none/moe · 1× mamba/none · 1× full/none · 1× none/moe · 1× mamba/none · 1× none/moe · 1× mamba/none · 1× none/moe · 1× mamba/none · 1× full/none · 1× none/moe · 1× mamba/none · 1× none/moe · 1× mamba/none · 1× none/moe · 1× mamba/none · 1× none/moe · 1× mamba/none · 1× full/none · 1× none/moe · 1× mamba/none · 1× none/moe · 1× mamba/none · 1× none/moe · 1× mamba/none · 1× none/moe · 1× mamba/none · 1× none/moeconfig.jsonconfig.jsonlayers_block_type
norms_per_layer1codemodelling codetransformers 5.18.0 nemotron_h: one pre-norm per block (each block is a mixer or an MLP)
mtp_layers1config.jsonconfig.jsonnum_nextn_predict_layers
mtp_layer{"mixer":"full","ffn":"moe","n":1}codemodelling codeMTP: attention + MoE blocks

Sources

Listed in the LLM Architecture Gallery checklist as “Nemotron 3.5 Lightning (30B-A3B)” (name only; see about).