llm-architectures-explained

/models

Nemotron 3 Nano 4B

NVIDIA · Nemotron 3 · open weights

Facts and where they come from

Released2026-03config.jsonconfig.jsonHugging Face repository creation date (api.createdAt)
Licenceotherconfig.jsonconfig.jsonREADME metadata: license
Total parameters4Blabmodel cardmodel name NVIDIA-Nemotron-3-Nano-4B
Active parametersnot disclosednot disclosed
Context length256K tokensconfig.jsonconfig.jsonmax_position_embeddings
Norm placementprecodemodelling codetransformers 5.18.0 nemotron_h: one norm before each mixer or MLP block
Norm typeRMSNormcodemodelling codetransformers 5.18.0 nemotron_h: one norm before each mixer or MLP block
QK-normnocodemodelling codeno q/k normalisation in the attention block
Positional encodingRoPEcodemodelling coderotary on the full head (default)
Parallel attention and MLPnocodemodelling codetransformers 5.18.0 nemotron_h: one norm before each mixer or MLP block

Architecture, drawn from the data

21× Mamba-2 + 4× GQA 40q/8kv. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

Nemotron 3 Nano 4B: layer stack and blockslayers (42)mixer / FFNlayer 0: Mamba-2: 96 heads × 80 · state 128layer 0: no FFNlayer 1: no mixerlayer 1: MLP: 12544layer 2: Mamba-2: 96 heads × 80 · state 128layer 2: no FFNlayer 3: no mixerlayer 3: MLP: 12544layer 4: Mamba-2: 96 heads × 80 · state 128layer 4: no FFNlayer 5: no mixerlayer 5: MLP: 12544layer 6: Mamba-2: 96 heads × 80 · state 128layer 6: no FFNlayer 7: Mamba-2: 96 heads × 80 · state 128layer 7: no FFNlayer 8: no mixerlayer 8: MLP: 12544layer 9: Mamba-2: 96 heads × 80 · state 128layer 9: no FFNlayer 10: no mixerlayer 10: MLP: 12544layer 11: Mamba-2: 96 heads × 80 · state 128layer 11: no FFNlayer 12: GQA: 40 query / 8 KV heads · head 128layer 12: no FFNlayer 13: no mixerlayer 13: MLP: 12544layer 14: Mamba-2: 96 heads × 80 · state 128layer 14: no FFNlayer 15: no mixerlayer 15: MLP: 12544layer 16: Mamba-2: 96 heads × 80 · state 128layer 16: no FFNlayer 17: GQA: 40 query / 8 KV heads · head 128layer 17: no FFNlayer 18: no mixerlayer 18: MLP: 12544layer 19: Mamba-2: 96 heads × 80 · state 128layer 19: no FFNlayer 20: no mixerlayer 20: MLP: 12544layer 21: Mamba-2: 96 heads × 80 · state 128layer 21: no FFNlayer 22: no mixerlayer 22: MLP: 12544layer 23: Mamba-2: 96 heads × 80 · state 128layer 23: no FFNlayer 24: GQA: 40 query / 8 KV heads · head 128layer 24: no FFNlayer 25: no mixerlayer 25: MLP: 12544layer 26: Mamba-2: 96 heads × 80 · state 128layer 26: no FFNlayer 27: no mixerlayer 27: MLP: 12544layer 28: Mamba-2: 96 heads × 80 · state 128layer 28: no FFNlayer 29: no mixerlayer 29: MLP: 12544layer 30: Mamba-2: 96 heads × 80 · state 128layer 30: no FFNlayer 31: Mamba-2: 96 heads × 80 · state 128layer 31: no FFNlayer 32: GQA: 40 query / 8 KV heads · head 128layer 32: no FFNlayer 33: no mixerlayer 33: MLP: 12544layer 34: Mamba-2: 96 heads × 80 · state 128layer 34: no FFNlayer 35: Mamba-2: 96 heads × 80 · state 128layer 35: no FFNlayer 36: Mamba-2: 96 heads × 80 · state 128layer 36: no FFNlayer 37: no mixerlayer 37: MLP: 12544layer 38: Mamba-2: 96 heads × 80 · state 128layer 38: no FFNlayer 39: no mixerlayer 39: MLP: 12544layer 40: Mamba-2: 96 heads × 80 · state 128layer 40: no FFNlayer 41: no mixerlayer 41: MLP: 1254402141× 21normMamba-2: 96 heads × 80 · state 128+× 17normMLP: 12544+× 4normGQA: 40 query / 8 KV heads · head 128+Mamba-2full attentiondense FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)3.97B
Active per token (modelled)3.97B
Without embeddings and output head3.15B total, 3.15B active
Published weights (Hugging Face count)3.97B
KV cache per token, BF16 (layers that grow with context)16 KiB
KV cache + state at 256K tokens, BF164.04 GiB
Decode FLOPs per token at 4K context7.58 GFLOP
Prefill FLOPs for a 4K prompt27 TFLOP

KV cache against context

Nemotron 3 Nano 4B: KV cache bytes against context length101001,00010,000100,00095 MiB950 MiBcontext (tokens)KV cache + state (BF16)Nemotron 3 Nano 4B

Compare with other models →

Every architecture field

FieldValueSource
d_model3,136config.jsonconfig.jsonhidden_size
vocab131,072config.jsonconfig.jsonvocab_size
tied_embeddingsfalseconfig.jsonconfig.jsontie_word_embeddings
mixers.full.typeattncodemodelling codeattention block
mixers.full.heads40config.jsonconfig.jsonnum_attention_heads
mixers.full.kv_heads8config.jsonconfig.jsonnum_key_value_heads
mixers.full.head_dim128config.jsonconfig.jsonhead_dim
mixers.mamba.typemamba2codemodelling codeMamba-2 mixer
mixers.mamba.heads96config.jsonconfig.jsonmamba_num_heads
mixers.mamba.head_dim80config.jsonconfig.jsonmamba_head_dim
mixers.mamba.state128config.jsonconfig.jsonssm_state_size
mixers.mamba.groups8config.jsonconfig.jsonn_groups
mixers.mamba.conv_kernel4config.jsonconfig.jsonconv_kernel
ffns.dense.typedensecodemodelling codeMLP block
ffns.dense.d_ff12,544config.jsonconfig.jsonintermediate_size
ffns.dense.gatedfalsecodemodelling codeMLP: two matrices, no gate
layout1× mamba/none · 1× none/dense · 1× mamba/none · 1× none/dense · 1× mamba/none · 1× none/dense · 2× mamba/none · 1× none/dense · 1× mamba/none · 1× none/dense · 1× mamba/none · 1× full/none · 1× none/dense · 1× mamba/none · 1× none/dense · 1× mamba/none · 1× full/none · 1× none/dense · 1× mamba/none · 1× none/dense · 1× mamba/none · 1× none/dense · 1× mamba/none · 1× full/none · 1× none/dense · 1× mamba/none · 1× none/dense · 1× mamba/none · 1× none/dense · 2× mamba/none · 1× full/none · 1× none/dense · 3× mamba/none · 1× none/dense · 1× mamba/none · 1× none/dense · 1× mamba/none · 1× none/denseconfig.jsonconfig.jsonhybrid_override_pattern
norms_per_layer1codemodelling codetransformers 5.18.0 nemotron_h: one pre-norm per block (each block is a mixer or an MLP)

Sources

Listed in the LLM Architecture Gallery checklist as “Nemotron 3 Nano (4B)” (name only; see about).