llm-architectures-explained

/models

Nanbeige4.2 3B

BOSS Zhipin · Nanbeige · open weights

Facts and where they come from

Released2026-07config.jsonconfig.jsonHugging Face repository creation date (api.createdAt)
Licenceapache-2.0config.jsonconfig.jsonREADME metadata: license
Total parameters4Blabmodel cardREADME comparison table: Total Params 4B
Active parametersnot disclosednot disclosed
Context length256K tokenslabmodel cardREADME: up to 262,144 tokens
Norm placementprecodemodelling codetransformers 5.18.0 / repo modelling code for nanbeige: input_layernorm before attention, post_attention_layernorm before the MLP
Norm typeRMSNormcodemodelling codetransformers 5.18.0 / repo modelling code for nanbeige: input_layernorm before attention, post_attention_layernorm before the MLP
QK-normyescodemodelling coderepo modeling_nanbeige.py: q_norm and k_norm
Positional encodingRoPEcodemodelling coderotary on the full head (default)
Parallel attention and MLPnocodemodelling codetransformers 5.18.0 / repo modelling code for nanbeige: input_layernorm before attention, post_attention_layernorm before the MLP

Architecture, drawn from the data

GQA 48q/8kv. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

Nanbeige4.2 3B: layer stack and blockslayers (22)mixer / FFNlayer 0: GQA: 48 query / 8 KV heads · head 128layer 0: Gated MLP: 10752layer 1: GQA: 48 query / 8 KV heads · head 128layer 1: Gated MLP: 10752layer 2: GQA: 48 query / 8 KV heads · head 128layer 2: Gated MLP: 10752layer 3: GQA: 48 query / 8 KV heads · head 128layer 3: Gated MLP: 10752layer 4: GQA: 48 query / 8 KV heads · head 128layer 4: Gated MLP: 10752layer 5: GQA: 48 query / 8 KV heads · head 128layer 5: Gated MLP: 10752layer 6: GQA: 48 query / 8 KV heads · head 128layer 6: Gated MLP: 10752layer 7: GQA: 48 query / 8 KV heads · head 128layer 7: Gated MLP: 10752layer 8: GQA: 48 query / 8 KV heads · head 128layer 8: Gated MLP: 10752layer 9: GQA: 48 query / 8 KV heads · head 128layer 9: Gated MLP: 10752layer 10: GQA: 48 query / 8 KV heads · head 128layer 10: Gated MLP: 10752layer 11: GQA: 48 query / 8 KV heads · head 128layer 11: Gated MLP: 10752layer 12: GQA: 48 query / 8 KV heads · head 128layer 12: Gated MLP: 10752layer 13: GQA: 48 query / 8 KV heads · head 128layer 13: Gated MLP: 10752layer 14: GQA: 48 query / 8 KV heads · head 128layer 14: Gated MLP: 10752layer 15: GQA: 48 query / 8 KV heads · head 128layer 15: Gated MLP: 10752layer 16: GQA: 48 query / 8 KV heads · head 128layer 16: Gated MLP: 10752layer 17: GQA: 48 query / 8 KV heads · head 128layer 17: Gated MLP: 10752layer 18: GQA: 48 query / 8 KV heads · head 128layer 18: Gated MLP: 10752layer 19: GQA: 48 query / 8 KV heads · head 128layer 19: Gated MLP: 10752layer 20: GQA: 48 query / 8 KV heads · head 128layer 20: Gated MLP: 10752layer 21: GQA: 48 query / 8 KV heads · head 128layer 21: Gated MLP: 1075201121× 22normGQA: 48 query / 8 KV heads · head 128+normGated MLP: 10752+full attentiondense FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)4.17B
Active per token (modelled)4.17B
Without embeddings and output head3.15B total, 3.15B active
Published weights (Hugging Face count)4.17B
KV cache per token, BF16 (layers that grow with context)88 KiB
KV cache + state at 256K tokens, BF1622 GiB
Decode FLOPs per token at 4K context18 GFLOP
Prefill FLOPs for a 4K prompt60.7 TFLOP

KV cache against context

Nanbeige4.2 3B: KV cache bytes against context length101001,00010,000100,000980 KiB9.5 MiB95 MiB950 MiB9.3 GiBcontext (tokens)KV cache + state (BF16)Nanbeige4.2 3B

Compare with other models →

Every architecture field

FieldValueSource
d_model3,072config.jsonconfig.jsonhidden_size
vocab166,144config.jsonconfig.jsonvocab_size
tied_embeddingsfalseconfig.jsonconfig.jsontie_word_embeddings
mixers.full.typeattncodemodelling codeattention block
mixers.full.heads48config.jsonconfig.jsonnum_attention_heads
mixers.full.kv_heads8config.jsonconfig.jsonnum_key_value_heads
mixers.full.head_dim128config.jsonconfig.jsonhead_dim
ffns.dense.typedensecodemodelling codeMLP block
ffns.dense.d_ff10,752config.jsonconfig.jsonintermediate_size
ffns.dense.gatedtruecodemodelling codeMLP: gated (SwiGLU/GeGLU)
layout22× full/denseconfig.jsonconfig.jsonnum_hidden_layers
loops2config.jsonconfig.jsonnum_loops

Sources

Listed in the LLM Architecture Gallery checklist as “Nanbeige 4.2 (3B)” (name only; see about).