llm-architectures-explained

/models

Ling 3.0 Flash

InclusionAI · Ling · open weights

Facts and where they come from

Released2026-08config.jsonconfig.jsonHugging Face repository creation date (api.createdAt)
Licencemitconfig.jsonconfig.jsonREADME metadata: license
Total parametersnot disclosednot disclosed
Active parametersnot disclosednot disclosed
Context length256K tokensconfig.jsonconfig.jsonmax_position_embeddings
Norm placementprecodemodelling codetransformers 5.18.0 / repo modelling code for bailing_hybrid: input_layernorm before attention, post_attention_layernorm before the MLP
Norm typeRMSNormcodemodelling codetransformers 5.18.0 / repo modelling code for bailing_hybrid: input_layernorm before attention, post_attention_layernorm before the MLP
QK-normyesconfig.jsonconfig.jsonuse_qk_norm
Positional encodingRoPE on 33.3% of each headconfig.jsonconfig.jsonqk_rope_head_dim / (qk_nope_head_dim + qk_rope_head_dim): decoupled RoPE
Parallel attention and MLPnocodemodelling codetransformers 5.18.0 / repo modelling code for bailing_hybrid: input_layernorm before attention, post_attention_layernorm before the MLP

Architecture, drawn from the data

35× Kimi Delta Attention + 7× MLA 32h, latent 512. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

Ling 3.0 Flash: layer stack and blockslayers (42)mixer / FFNlayer 0: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 0: Gated MLP: 6144layer 1: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 1: Gated MLP: 6144layer 2: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 2: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 3: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 3: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 4: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 4: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 5: MLA: 32 heads · KV latent 512 + RoPE 64layer 5: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 6: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 6: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 7: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 7: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 8: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 8: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 9: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 9: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 10: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 10: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 11: MLA: 32 heads · KV latent 512 + RoPE 64layer 11: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 12: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 12: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 13: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 13: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 14: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 14: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 15: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 15: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 16: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 16: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 17: MLA: 32 heads · KV latent 512 + RoPE 64layer 17: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 18: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 18: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 19: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 19: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 20: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 20: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 21: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 21: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 22: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 22: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 23: MLA: 32 heads · KV latent 512 + RoPE 64layer 23: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 24: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 24: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 25: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 25: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 26: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 26: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 27: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 27: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 28: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 28: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 29: MLA: 32 heads · KV latent 512 + RoPE 64layer 29: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 30: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 30: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 31: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 31: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 32: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 32: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 33: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 33: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 34: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 34: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 35: MLA: 32 heads · KV latent 512 + RoPE 64layer 35: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 36: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 36: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 37: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 37: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 38: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 38: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 39: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 39: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 40: Kimi Delta Attention: 32 heads · state 128×128 per headlayer 40: MoE: 512 experts, 8 active · expert 768 · 1 sharedlayer 41: MLA: 32 heads · KV latent 512 + RoPE 64layer 41: MoE: 512 experts, 8 active · expert 768 · 1 shared02141× 33normKimi Delta Attention: 32 heads · state 128×128 per head+normMoE: 512 experts, 8 active · expert 768 · 1 shared+× 7normMLA: 32 heads · KV latent 512 + RoPE 64+normMoE: 512 experts, 8 active · expert 768 · 1 shared+× 2normKimi Delta Attention: 32 heads · state 128×128 per head+normGated MLP: 6144+Kimi Delta AttentionMLAdense FFNMoE FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)124B
Active per token (modelled)5.2B
Without embeddings and output head123B total, 4.39B active
Multi-token-prediction layers (extra)3.07B
Published weights (Hugging Face count)127B
KV cache per token, BF16 (layers that grow with context)7.88 KiB
KV cache + state at 256K tokens, BF162.01 GiB
Decode FLOPs per token at 4K context10.3 GFLOP
Prefill FLOPs for a 4K prompt37.6 TFLOP

KV cache against context

Ling 3.0 Flash: KV cache bytes against context length101001,00010,000100,00095 MiB950 MiBcontext (tokens)KV cache + state (BF16)Ling 3.0 Flash

Compare with other models →

Every architecture field

FieldValueSource
d_model2,560config.jsonconfig.jsonhidden_size
vocab157,184config.jsonconfig.jsonvocab_size
tied_embeddingsfalseconfig.jsonconfig.jsontie_word_embeddings
mixers.mla.typemlacodemodelling codemulti-head latent attention
mixers.mla.heads32config.jsonconfig.jsonnum_attention_heads
mixers.mla.q_lora_ranknullcodemodelling codeq_lora_rank: null (full-rank query)
mixers.mla.kv_lora_rank512config.jsonconfig.jsonkv_lora_rank
mixers.mla.qk_nope128config.jsonconfig.jsonqk_nope_head_dim
mixers.mla.qk_rope64config.jsonconfig.jsonqk_rope_head_dim
mixers.mla.v_head_dim128config.jsonconfig.jsonv_head_dim
mixers.linear.typekdacodemodelling codeKimi Delta Attention (BailingMoeV3KimiDeltaAttention)
mixers.linear.k_heads32config.jsonconfig.jsonnum_attention_heads
mixers.linear.v_heads32config.jsonconfig.jsonnum_attention_heads
mixers.linear.k_head_dim128config.jsonconfig.jsonhead_dim
mixers.linear.v_head_dim128config.jsonconfig.jsonhead_dim
mixers.linear.conv_kernel4config.jsonconfig.jsonshort_conv_kernel_size
ffns.dense.typedensecodemodelling codeMLP block
ffns.dense.d_ff6,144config.jsonconfig.jsonintermediate_size
ffns.dense.gatedtruecodemodelling codeMLP: gated (SwiGLU/GeGLU)
ffns.moe.typemoecodemodelling codeMoE block
ffns.moe.experts512config.jsonconfig.jsonnum_experts
ffns.moe.active8config.jsonconfig.jsonnum_experts_per_tok
ffns.moe.d_expert768config.jsonconfig.jsonmoe_intermediate_size
ffns.moe.gatedtruecodemodelling codeexperts are gated MLPs
ffns.moe.shared1config.jsonconfig.jsonnum_shared_experts
ffns.moe.d_shared768config.jsonconfig.jsonmoe_shared_expert_intermediate_size
layout2× linear/dense · 3× linear/moe · 1× mla/moe · 5× linear/moe · 1× mla/moe · 5× linear/moe · 1× mla/moe · 5× linear/moe · 1× mla/moe · 5× linear/moe · 1× mla/moe · 5× linear/moe · 1× mla/moe · 5× linear/moe · 1× mla/moecodemodelling codemodeling_bailing_moe_v3.py: softmax attention when (layer_idx + 1) % layer_group_size == 0 or in the tail
mtp_layers1config.jsonconfig.jsonnum_nextn_predict_layers
mtp_layer{"mixer":"mla","ffn":"moe","n":1}codemodelling codeMTP layer: MLA + MoE

Sources

Listed in the LLM Architecture Gallery checklist as “Ling 3.0 Flash (124B)” (name only; see about).