llm-architectures-explained

/models

Ling 2.6 1T

InclusionAI · Ling · open weights

Facts and where they come from

Released2026-04config.jsonconfig.jsonHugging Face repository creation date (api.createdAt)
Licencemitconfig.jsonconfig.jsonREADME metadata: license
Total parameters1Tlabmodel cardmodel name Ling-2.6-1T
Active parametersnot disclosednot disclosed
Context length256K tokensconfig.jsonconfig.jsonmax_position_embeddings
Norm placementprecodemodelling codetransformers 5.18.0 / repo modelling code for bailing_hybrid: input_layernorm before attention, post_attention_layernorm before the MLP
Norm typeRMSNormcodemodelling codetransformers 5.18.0 / repo modelling code for bailing_hybrid: input_layernorm before attention, post_attention_layernorm before the MLP
QK-normyesconfig.jsonconfig.jsonuse_qk_norm
Positional encodingRoPE on 33.3% of each headconfig.jsonconfig.jsonqk_rope_head_dim / (qk_nope_head_dim + qk_rope_head_dim): decoupled RoPE
Parallel attention and MLPnocodemodelling codetransformers 5.18.0 / repo modelling code for bailing_hybrid: input_layernorm before attention, post_attention_layernorm before the MLP

Architecture, drawn from the data

70× Linear attention + 10× MLA 64h, latent 512. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

Ling 2.6 1T: layer stack and blockslayers (80)mixer / FFNlayer 0: Linear attention: 64 heads × 128layer 0: Gated MLP: 18432layer 1: Linear attention: 64 heads × 128layer 1: Gated MLP: 18432layer 2: Linear attention: 64 heads × 128layer 2: Gated MLP: 18432layer 3: Linear attention: 64 heads × 128layer 3: Gated MLP: 18432layer 4: Linear attention: 64 heads × 128layer 4: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 5: Linear attention: 64 heads × 128layer 5: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 6: Linear attention: 64 heads × 128layer 6: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 7: MLA: 64 heads · KV latent 512 + RoPE 64 · Q latent 1536layer 7: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 8: Linear attention: 64 heads × 128layer 8: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 9: Linear attention: 64 heads × 128layer 9: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 10: Linear attention: 64 heads × 128layer 10: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 11: Linear attention: 64 heads × 128layer 11: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 12: Linear attention: 64 heads × 128layer 12: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 13: Linear attention: 64 heads × 128layer 13: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 14: Linear attention: 64 heads × 128layer 14: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 15: MLA: 64 heads · KV latent 512 + RoPE 64 · Q latent 1536layer 15: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 16: Linear attention: 64 heads × 128layer 16: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 17: Linear attention: 64 heads × 128layer 17: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 18: Linear attention: 64 heads × 128layer 18: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 19: Linear attention: 64 heads × 128layer 19: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 20: Linear attention: 64 heads × 128layer 20: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 21: Linear attention: 64 heads × 128layer 21: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 22: Linear attention: 64 heads × 128layer 22: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 23: MLA: 64 heads · KV latent 512 + RoPE 64 · Q latent 1536layer 23: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 24: Linear attention: 64 heads × 128layer 24: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 25: Linear attention: 64 heads × 128layer 25: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 26: Linear attention: 64 heads × 128layer 26: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 27: Linear attention: 64 heads × 128layer 27: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 28: Linear attention: 64 heads × 128layer 28: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 29: Linear attention: 64 heads × 128layer 29: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 30: Linear attention: 64 heads × 128layer 30: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 31: MLA: 64 heads · KV latent 512 + RoPE 64 · Q latent 1536layer 31: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 32: Linear attention: 64 heads × 128layer 32: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 33: Linear attention: 64 heads × 128layer 33: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 34: Linear attention: 64 heads × 128layer 34: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 35: Linear attention: 64 heads × 128layer 35: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 36: Linear attention: 64 heads × 128layer 36: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 37: Linear attention: 64 heads × 128layer 37: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 38: Linear attention: 64 heads × 128layer 38: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 39: MLA: 64 heads · KV latent 512 + RoPE 64 · Q latent 1536layer 39: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 40: Linear attention: 64 heads × 128layer 40: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 41: Linear attention: 64 heads × 128layer 41: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 42: Linear attention: 64 heads × 128layer 42: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 43: Linear attention: 64 heads × 128layer 43: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 44: Linear attention: 64 heads × 128layer 44: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 45: Linear attention: 64 heads × 128layer 45: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 46: Linear attention: 64 heads × 128layer 46: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 47: MLA: 64 heads · KV latent 512 + RoPE 64 · Q latent 1536layer 47: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 48: Linear attention: 64 heads × 128layer 48: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 49: Linear attention: 64 heads × 128layer 49: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 50: Linear attention: 64 heads × 128layer 50: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 51: Linear attention: 64 heads × 128layer 51: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 52: Linear attention: 64 heads × 128layer 52: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 53: Linear attention: 64 heads × 128layer 53: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 54: Linear attention: 64 heads × 128layer 54: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 55: MLA: 64 heads · KV latent 512 + RoPE 64 · Q latent 1536layer 55: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 56: Linear attention: 64 heads × 128layer 56: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 57: Linear attention: 64 heads × 128layer 57: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 58: Linear attention: 64 heads × 128layer 58: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 59: Linear attention: 64 heads × 128layer 59: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 60: Linear attention: 64 heads × 128layer 60: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 61: Linear attention: 64 heads × 128layer 61: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 62: Linear attention: 64 heads × 128layer 62: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 63: MLA: 64 heads · KV latent 512 + RoPE 64 · Q latent 1536layer 63: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 64: Linear attention: 64 heads × 128layer 64: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 65: Linear attention: 64 heads × 128layer 65: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 66: Linear attention: 64 heads × 128layer 66: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 67: Linear attention: 64 heads × 128layer 67: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 68: Linear attention: 64 heads × 128layer 68: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 69: Linear attention: 64 heads × 128layer 69: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 70: Linear attention: 64 heads × 128layer 70: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 71: MLA: 64 heads · KV latent 512 + RoPE 64 · Q latent 1536layer 71: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 72: Linear attention: 64 heads × 128layer 72: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 73: Linear attention: 64 heads × 128layer 73: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 74: Linear attention: 64 heads × 128layer 74: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 75: Linear attention: 64 heads × 128layer 75: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 76: Linear attention: 64 heads × 128layer 76: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 77: Linear attention: 64 heads × 128layer 77: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 78: Linear attention: 64 heads × 128layer 78: MoE: 256 experts, 8 active · expert 2048 · 1 sharedlayer 79: MLA: 64 heads · KV latent 512 + RoPE 64 · Q latent 1536layer 79: MoE: 256 experts, 8 active · expert 2048 · 1 shared04079× 66normLinear attention: 64 heads × 128+normMoE: 256 experts, 8 active · expert 2048 · 1 shared+× 10normMLA: 64 heads · KV latent 512 + RoPE 64 · Q latent 1536+normMoE: 256 experts, 8 active · expert 2048 · 1 shared+× 4normLinear attention: 64 heads × 128+normGated MLP: 18432+linear attentionMLAdense FFNMoE FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)1.01T
Active per token (modelled)63.6B
Without embeddings and output head1.01T total, 61B active
Multi-token-prediction layers (extra)13.2B
Published weights (Hugging Face count)1.03T
KV cache per token, BF16 (layers that grow with context)11.3 KiB
KV cache + state at 256K tokens, BF162.95 GiB
Decode FLOPs per token at 4K context127 GFLOP
Prefill FLOPs for a 4K prompt504 TFLOP

KV cache against context

Ling 2.6 1T: KV cache bytes against context length101001,00010,000100,000950 MiBcontext (tokens)KV cache + state (BF16)Ling 2.6 1T

Compare with other models →

Every architecture field

FieldValueSource
d_model8,192config.jsonconfig.jsonhidden_size
vocab157,184config.jsonconfig.jsonvocab_size
tied_embeddingsfalseconfig.jsonconfig.jsontie_word_embeddings
mixers.mla.typemlacodemodelling codemulti-head latent attention
mixers.mla.heads64config.jsonconfig.jsonnum_attention_heads
mixers.mla.q_lora_rank1,536config.jsonconfig.jsonq_lora_rank
mixers.mla.kv_lora_rank512config.jsonconfig.jsonkv_lora_rank
mixers.mla.qk_nope128config.jsonconfig.jsonqk_nope_head_dim
mixers.mla.qk_rope64config.jsonconfig.jsonqk_rope_head_dim
mixers.mla.v_head_dim128config.jsonconfig.jsonv_head_dim
mixers.linear.typelinearcodemodelling codeLightning-style linear attention (BailingMoeV2_5LinearAttention)
mixers.linear.heads64config.jsonconfig.jsonnum_kv_heads_for_linear_attn
mixers.linear.head_dim128config.jsonconfig.jsonhead_dim
ffns.dense.typedensecodemodelling codeMLP block
ffns.dense.d_ff18,432config.jsonconfig.jsonintermediate_size
ffns.dense.gatedtruecodemodelling codeMLP: gated (SwiGLU/GeGLU)
ffns.moe.typemoecodemodelling codeMoE block
ffns.moe.experts256config.jsonconfig.jsonnum_experts
ffns.moe.active8config.jsonconfig.jsonnum_experts_per_tok
ffns.moe.d_expert2,048config.jsonconfig.jsonmoe_intermediate_size
ffns.moe.gatedtruecodemodelling codeexperts are gated MLPs
ffns.moe.shared1config.jsonconfig.jsonnum_shared_experts
ffns.moe.d_shared2,048config.jsonconfig.jsonmoe_shared_expert_intermediate_size
layout4× linear/dense · 3× linear/moe · 1× mla/moe · 7× linear/moe · 1× mla/moe · 7× linear/moe · 1× mla/moe · 7× linear/moe · 1× mla/moe · 7× linear/moe · 1× mla/moe · 7× linear/moe · 1× mla/moe · 7× linear/moe · 1× mla/moe · 7× linear/moe · 1× mla/moe · 7× linear/moe · 1× mla/moe · 7× linear/moe · 1× mla/moecodemodelling codemodeling_bailing_moe_v2_5.py: softmax attention when (layer_idx + 1) % layer_group_size == 0 or in the tail
mtp_layers1config.jsonconfig.jsonnum_nextn_predict_layers
mtp_layer{"mixer":"mla","ffn":"moe","n":1}codemodelling codeMTP layer: MLA + MoE

Sources

Listed in the LLM Architecture Gallery checklist as “Ling 2.6 (1T)” (name only; see about).