llm-architectures-explained

/models

Kimi Linear 48B-A3B

Moonshot AI · Kimi Linear · open weights

Facts and where they come from

Released2025-10config.jsonconfig.jsonHugging Face repository creation date (api.createdAt)
Licencemitconfig.jsonconfig.jsonREADME metadata: license
Total parameters48Blabmodel cardmodel name Kimi-Linear-48B-A3B
Active parameters3Blabmodel cardmodel name ...-A3B
Context lengthnot disclosednot disclosed
Norm placementprecodemodelling codetransformers 5.18.0 / repo modelling code for kimi_linear: input_layernorm before attention, post_attention_layernorm before the MLP
Norm typeRMSNormcodemodelling codetransformers 5.18.0 / repo modelling code for kimi_linear: input_layernorm before attention, post_attention_layernorm before the MLP
QK-normnocodemodelling codeno q/k normalisation in the attention block (MLA normalises its latent vectors, which is not QK-norm)
Positional encodingRoPE on 33.3% of each head; MLA layers use no RoPE; the linear-attention layers carry orderconfig.jsonconfig.jsonmla_use_nope
Parallel attention and MLPnocodemodelling codetransformers 5.18.0 / repo modelling code for kimi_linear: input_layernorm before attention, post_attention_layernorm before the MLP

Architecture, drawn from the data

20× Kimi Delta Attention + 7× MLA 32h, latent 512. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

Kimi Linear 48B-A3B: layer stack and blockslayers (27)mixer / FFNlayer 0: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 0: Gated MLP: 9216layer 1: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 1: MoE: 256 experts, 8 active · expert 1024 · 1 sharedlayer 2: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 2: MoE: 256 experts, 8 active · expert 1024 · 1 sharedlayer 3: MLA: 32 heads · KV latent 512 + RoPE 64layer 3: MoE: 256 experts, 8 active · expert 1024 · 1 sharedlayer 4: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 4: MoE: 256 experts, 8 active · expert 1024 · 1 sharedlayer 5: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 5: MoE: 256 experts, 8 active · expert 1024 · 1 sharedlayer 6: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 6: MoE: 256 experts, 8 active · expert 1024 · 1 sharedlayer 7: MLA: 32 heads · KV latent 512 + RoPE 64layer 7: MoE: 256 experts, 8 active · expert 1024 · 1 sharedlayer 8: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 8: MoE: 256 experts, 8 active · expert 1024 · 1 sharedlayer 9: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 9: MoE: 256 experts, 8 active · expert 1024 · 1 sharedlayer 10: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 10: MoE: 256 experts, 8 active · expert 1024 · 1 sharedlayer 11: MLA: 32 heads · KV latent 512 + RoPE 64layer 11: MoE: 256 experts, 8 active · expert 1024 · 1 sharedlayer 12: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 12: MoE: 256 experts, 8 active · expert 1024 · 1 sharedlayer 13: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 13: MoE: 256 experts, 8 active · expert 1024 · 1 sharedlayer 14: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 14: MoE: 256 experts, 8 active · expert 1024 · 1 sharedlayer 15: MLA: 32 heads · KV latent 512 + RoPE 64layer 15: MoE: 256 experts, 8 active · expert 1024 · 1 sharedlayer 16: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 16: MoE: 256 experts, 8 active · expert 1024 · 1 sharedlayer 17: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 17: MoE: 256 experts, 8 active · expert 1024 · 1 sharedlayer 18: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 18: MoE: 256 experts, 8 active · expert 1024 · 1 sharedlayer 19: MLA: 32 heads · KV latent 512 + RoPE 64layer 19: MoE: 256 experts, 8 active · expert 1024 · 1 sharedlayer 20: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 20: MoE: 256 experts, 8 active · expert 1024 · 1 sharedlayer 21: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 21: MoE: 256 experts, 8 active · expert 1024 · 1 sharedlayer 22: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 22: MoE: 256 experts, 8 active · expert 1024 · 1 sharedlayer 23: MLA: 32 heads · KV latent 512 + RoPE 64layer 23: MoE: 256 experts, 8 active · expert 1024 · 1 sharedlayer 24: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 24: MoE: 256 experts, 8 active · expert 1024 · 1 sharedlayer 25: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 25: MoE: 256 experts, 8 active · expert 1024 · 1 sharedlayer 26: MLA: 32 heads · KV latent 512 + RoPE 64layer 26: MoE: 256 experts, 8 active · expert 1024 · 1 shared01326× 19normKimi Delta Attention: 32 heads · state 128×128 per head+normMoE: 256 experts, 8 active · expert 1024 · 1 shared+× 7normMLA: 32 heads · KV latent 512 + RoPE 64+normMoE: 256 experts, 8 active · expert 1024 · 1 shared+× 1normKimi Delta Attention: 32 heads · state 128×128 per head+normGated MLP: 9216+Kimi Delta AttentionMLAdense FFNMoE FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)49.3B
Active per token (modelled)3.67B
Without embeddings and output head48.6B total, 2.92B active
Published weights (Hugging Face count)49.1B
KV cache per token, BF16 (layers that grow with context)7.88 KiB
KV cache + state at 128K tokens, BF161.01 GiB
Decode FLOPs per token at 4K context7.24 GFLOP
Prefill FLOPs for a 4K prompt25.4 TFLOP

KV cache against context

Kimi Linear 48B-A3B: KV cache bytes against context length101001,00010,000100,00095 MiB950 MiBcontext (tokens)KV cache + state (BF16)Kimi Linear 48B-A3B

Compare with other models →

Every architecture field

FieldValueSource
d_model2,304config.jsonconfig.jsonhidden_size
vocab163,840config.jsonconfig.jsonvocab_size
tied_embeddingsfalseconfig.jsonconfig.jsontie_word_embeddings
mixers.mla.typemlacodemodelling codemulti-head latent attention
mixers.mla.heads32config.jsonconfig.jsonnum_attention_heads
mixers.mla.q_lora_ranknullcodemodelling codeq_lora_rank: null (full-rank query)
mixers.mla.kv_lora_rank512config.jsonconfig.jsonkv_lora_rank
mixers.mla.qk_nope128config.jsonconfig.jsonqk_nope_head_dim
mixers.mla.qk_rope64config.jsonconfig.jsonqk_rope_head_dim
mixers.mla.v_head_dim128config.jsonconfig.jsonv_head_dim
mixers.kda.typekdacodemodelling codeKimi Delta Attention
mixers.kda.k_heads32config.jsonconfig.jsonlinear_attn_config.num_heads
mixers.kda.v_heads32config.jsonconfig.jsonlinear_attn_config.num_heads
mixers.kda.k_head_dim128config.jsonconfig.jsonlinear_attn_config.head_dim
mixers.kda.v_head_dim128config.jsonconfig.jsonlinear_attn_config.head_dim
mixers.kda.conv_kernel4config.jsonconfig.jsonlinear_attn_config.short_conv_kernel_size
ffns.dense.typedensecodemodelling codeMLP block
ffns.dense.d_ff9,216config.jsonconfig.jsonintermediate_size
ffns.dense.gatedtruecodemodelling codeMLP: gated (SwiGLU/GeGLU)
ffns.moe.typemoecodemodelling codeMoE block
ffns.moe.experts256config.jsonconfig.jsonnum_experts
ffns.moe.active8config.jsonconfig.jsonnum_experts_per_token
ffns.moe.d_expert1,024config.jsonconfig.jsonmoe_intermediate_size
ffns.moe.gatedtruecodemodelling codeexperts are gated MLPs
ffns.moe.shared1config.jsonconfig.jsonnum_shared_experts
ffns.moe.d_shared1,024config.jsonconfig.jsonmoe_intermediate_size
layout1× kda/dense · 2× kda/moe · 1× mla/moe · 3× kda/moe · 1× mla/moe · 3× kda/moe · 1× mla/moe · 3× kda/moe · 1× mla/moe · 3× kda/moe · 1× mla/moe · 3× kda/moe · 1× mla/moe · 2× kda/moe · 1× mla/moeconfig.jsonconfig.jsonlinear_attn_config.full_attn_layers (1-based), first_k_dense_replace

Sources

Listed in the LLM Architecture Gallery checklist as “Kimi Linear (48B-A3B)” (name only; see about).