llm-architectures-explained

/models

Sarvam 30B

Sarvam AI · Sarvam · open weights

Facts and where they come from

Released2026-03config.jsonconfig.jsonHugging Face repository creation date (api.createdAt)
Licenceapache-2.0config.jsonconfig.jsonREADME metadata: license
Total parameters30Blabmodel cardmodel name sarvam-30b
Active parametersnot disclosednot disclosed
Context length128K tokensconfig.jsonconfig.jsonmax_position_embeddings
Norm placementprecodemodelling codetransformers 5.18.0 / repo modelling code for sarvam_moe: input_layernorm before attention, post_attention_layernorm before the MLP
Norm typeRMSNormcodemodelling codetransformers 5.18.0 / repo modelling code for sarvam_moe: input_layernorm before attention, post_attention_layernorm before the MLP
QK-normyesconfig.jsonconfig.jsonuse_qk_norm
Positional encodingRoPEcodemodelling coderotary on the full head (default)
Parallel attention and MLPnocodemodelling codetransformers 5.18.0 / repo modelling code for sarvam_moe: input_layernorm before attention, post_attention_layernorm before the MLP

Architecture, drawn from the data

GQA 64q/4kv. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

Sarvam 30B: layer stack and blockslayers (19)mixer / FFNlayer 0: GQA: 64 query / 4 KV heads · head 64layer 0: Gated MLP: 8192layer 1: GQA: 64 query / 4 KV heads · head 64layer 1: MoE: 128 experts, 6 active · expert 1024 · 1 sharedlayer 2: GQA: 64 query / 4 KV heads · head 64layer 2: MoE: 128 experts, 6 active · expert 1024 · 1 sharedlayer 3: GQA: 64 query / 4 KV heads · head 64layer 3: MoE: 128 experts, 6 active · expert 1024 · 1 sharedlayer 4: GQA: 64 query / 4 KV heads · head 64layer 4: MoE: 128 experts, 6 active · expert 1024 · 1 sharedlayer 5: GQA: 64 query / 4 KV heads · head 64layer 5: MoE: 128 experts, 6 active · expert 1024 · 1 sharedlayer 6: GQA: 64 query / 4 KV heads · head 64layer 6: MoE: 128 experts, 6 active · expert 1024 · 1 sharedlayer 7: GQA: 64 query / 4 KV heads · head 64layer 7: MoE: 128 experts, 6 active · expert 1024 · 1 sharedlayer 8: GQA: 64 query / 4 KV heads · head 64layer 8: MoE: 128 experts, 6 active · expert 1024 · 1 sharedlayer 9: GQA: 64 query / 4 KV heads · head 64layer 9: MoE: 128 experts, 6 active · expert 1024 · 1 sharedlayer 10: GQA: 64 query / 4 KV heads · head 64layer 10: MoE: 128 experts, 6 active · expert 1024 · 1 sharedlayer 11: GQA: 64 query / 4 KV heads · head 64layer 11: MoE: 128 experts, 6 active · expert 1024 · 1 sharedlayer 12: GQA: 64 query / 4 KV heads · head 64layer 12: MoE: 128 experts, 6 active · expert 1024 · 1 sharedlayer 13: GQA: 64 query / 4 KV heads · head 64layer 13: MoE: 128 experts, 6 active · expert 1024 · 1 sharedlayer 14: GQA: 64 query / 4 KV heads · head 64layer 14: MoE: 128 experts, 6 active · expert 1024 · 1 sharedlayer 15: GQA: 64 query / 4 KV heads · head 64layer 15: MoE: 128 experts, 6 active · expert 1024 · 1 sharedlayer 16: GQA: 64 query / 4 KV heads · head 64layer 16: MoE: 128 experts, 6 active · expert 1024 · 1 sharedlayer 17: GQA: 64 query / 4 KV heads · head 64layer 17: MoE: 128 experts, 6 active · expert 1024 · 1 sharedlayer 18: GQA: 64 query / 4 KV heads · head 64layer 18: MoE: 128 experts, 6 active · expert 1024 · 1 shared0918× 18normGQA: 64 query / 4 KV heads · head 64+normMoE: 128 experts, 6 active · expert 1024 · 1 shared+× 1normGQA: 64 query / 4 KV heads · head 64+normGated MLP: 8192+full attentiondense FFNMoE FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)32.2B
Active per token (modelled)4.52B
Without embeddings and output head30B total, 2.37B active
Published weights (Hugging Face count)32.2B
KV cache per token, BF16 (layers that grow with context)19 KiB
KV cache + state at 128K tokens, BF162.38 GiB
Decode FLOPs per token at 4K context8.17 GFLOP
Prefill FLOPs for a 4K prompt22.1 TFLOP

KV cache against context

Sarvam 30B: KV cache bytes against context length101001,00010,000100,00098 KiB980 KiB9.5 MiB95 MiB950 MiBcontext (tokens)KV cache + state (BF16)Sarvam 30B

Compare with other models →

Every architecture field

FieldValueSource
d_model4,096config.jsonconfig.jsonhidden_size
vocab262,144config.jsonconfig.jsonvocab_size
tied_embeddingsfalseconfig.jsonconfig.jsontie_word_embeddings
mixers.full.typeattncodemodelling codeattention block
mixers.full.heads64config.jsonconfig.jsonnum_attention_heads
mixers.full.kv_heads4config.jsonconfig.jsonnum_key_value_heads
mixers.full.head_dim64config.jsonconfig.jsonhead_dim
mixers.full.qk_normtrueconfig.jsonconfig.jsonuse_qk_norm
ffns.dense.typedensecodemodelling codeMLP block
ffns.dense.d_ff8,192config.jsonconfig.jsonintermediate_size
ffns.dense.gatedtruecodemodelling codeMLP: gated (SwiGLU/GeGLU)
ffns.moe.typemoecodemodelling codeMoE block
ffns.moe.experts128config.jsonconfig.jsonnum_experts
ffns.moe.active6config.jsonconfig.jsonnum_experts_per_tok
ffns.moe.d_expert1,024config.jsonconfig.jsonmoe_intermediate_size
ffns.moe.gatedtruecodemodelling codeexperts are gated MLPs
ffns.moe.shared1config.jsonconfig.jsonnum_shared_experts
ffns.moe.d_shared1,024config.jsonconfig.jsonmoe_shared_expert_intermediate_size
layout1× full/dense · 18× full/moeconfig.jsonconfig.jsonnum_hidden_layers, first_k_dense_replace

Sources

Listed in the LLM Architecture Gallery checklist as “Sarvam (30B)” (name only; see about).