llm-architectures-explained

/models

Solar Open 2

Upstage · Solar · open weights

Facts and where they come from

Released2026-07config.jsonconfig.jsonHugging Face repository creation date (api.createdAt)
Licenceotherconfig.jsonconfig.jsonREADME metadata: license
Total parameters250Blabmodel cardREADME: Total Parameters 250B (250,287,794,944)
Active parameters15Blabmodel cardREADME: Active Parameters 15B
Context length1M tokenslabmodel cardREADME: Context Length 1M
Norm placementprecodemodelling codetransformers 5.18.0 / repo modelling code for solar_open2: input_layernorm before attention, post_attention_layernorm before the MLP
Norm typeRMSNormcodemodelling codetransformers 5.18.0 / repo modelling code for solar_open2: input_layernorm before attention, post_attention_layernorm before the MLP
QK-normnocodemodelling codeno q/k normalisation in the attention block
Positional encodingRoPE; attention layers use no RoPE; the linear-attention layers carry orderconfig.jsonconfig.jsonuse_rope: false
Parallel attention and MLPnocodemodelling codetransformers 5.18.0 / repo modelling code for solar_open2: input_layernorm before attention, post_attention_layernorm before the MLP

Architecture, drawn from the data

12× GQA 64q/8kv + 36× Kimi Delta Attention. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

Solar Open 2: layer stack and blockslayers (48)mixer / FFNlayer 0: GQA: 64 query / 8 KV heads · head 128 · elementwise output gatelayer 0: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 1: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 1: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 2: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 2: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 3: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 3: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 4: GQA: 64 query / 8 KV heads · head 128 · elementwise output gatelayer 4: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 5: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 5: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 6: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 6: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 7: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 7: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 8: GQA: 64 query / 8 KV heads · head 128 · elementwise output gatelayer 8: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 9: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 9: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 10: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 10: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 11: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 11: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 12: GQA: 64 query / 8 KV heads · head 128 · elementwise output gatelayer 12: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 13: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 13: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 14: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 14: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 15: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 15: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 16: GQA: 64 query / 8 KV heads · head 128 · elementwise output gatelayer 16: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 17: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 17: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 18: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 18: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 19: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 19: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 20: GQA: 64 query / 8 KV heads · head 128 · elementwise output gatelayer 20: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 21: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 21: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 22: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 22: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 23: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 23: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 24: GQA: 64 query / 8 KV heads · head 128 · elementwise output gatelayer 24: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 25: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 25: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 26: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 26: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 27: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 27: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 28: GQA: 64 query / 8 KV heads · head 128 · elementwise output gatelayer 28: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 29: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 29: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 30: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 30: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 31: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 31: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 32: GQA: 64 query / 8 KV heads · head 128 · elementwise output gatelayer 32: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 33: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 33: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 34: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 34: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 35: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 35: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 36: GQA: 64 query / 8 KV heads · head 128 · elementwise output gatelayer 36: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 37: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 37: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 38: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 38: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 39: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 39: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 40: GQA: 64 query / 8 KV heads · head 128 · elementwise output gatelayer 40: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 41: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 41: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 42: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 42: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 43: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 43: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 44: GQA: 64 query / 8 KV heads · head 128 · elementwise output gatelayer 44: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 45: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 45: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 46: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 46: MoE: 320 experts, 8 active · expert 1280 · 1 sharedlayer 47: Kimi Delta Attention: 64 heads · state 128×128 per headlayer 47: MoE: 320 experts, 8 active · expert 1280 · 1 shared02447× 36normKimi Delta Attention: 64 heads · state 128×128 per head+normMoE: 320 experts, 8 active · expert 1280 · 1 shared+× 12normGQA: 64 query / 8 KV heads · head 128 · elementwise output gate+normMoE: 320 experts, 8 active · expert 1280 · 1 shared+full attentionKimi Delta AttentionMoE FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)251B
Active per token (modelled)15.9B
Without embeddings and output head250B total, 14.3B active
Published weights (Hugging Face count)250B
KV cache per token, BF16 (layers that grow with context)48 KiB
KV cache + state at 1M tokens, BF1648.1 GiB
Decode FLOPs per token at 4K context32.1 GFLOP
Prefill FLOPs for a 4K prompt122 TFLOP

KV cache against context

Solar Open 2: KV cache bytes against context length101001,00010,000100,0001,000,00095 MiB950 MiB9.3 GiBcontext (tokens)KV cache + state (BF16)Solar Open 2

Compare with other models →

Every architecture field

FieldValueSource
d_model4,096config.jsonconfig.jsonhidden_size
vocab196,608config.jsonconfig.jsonvocab_size
tied_embeddingsfalseconfig.jsonconfig.jsontie_word_embeddings
mixers.full.typeattncodemodelling codeattention block
mixers.full.heads64config.jsonconfig.jsonnum_attention_heads
mixers.full.kv_heads8config.jsonconfig.jsonnum_key_value_heads
mixers.full.head_dim128config.jsonconfig.jsonhead_dim
mixers.full.gateelementwisecodemodelling codeuse_gqa_gate: true
mixers.kda.typekdacodemodelling codeKimi Delta Attention
mixers.kda.k_heads64config.jsonconfig.jsonlinear_attn_config.num_heads
mixers.kda.v_heads64config.jsonconfig.jsonlinear_attn_config.num_heads
mixers.kda.k_head_dim128config.jsonconfig.jsonlinear_attn_config.head_dim
mixers.kda.v_head_dim128config.jsonconfig.jsonlinear_attn_config.head_dim
mixers.kda.conv_kernel4config.jsonconfig.jsonlinear_attn_config.short_conv_kernel_size
ffns.moe.typemoecodemodelling codeMoE block
ffns.moe.experts320config.jsonconfig.jsonn_routed_experts
ffns.moe.active8config.jsonconfig.jsonnum_experts_per_tok
ffns.moe.d_expert1,280config.jsonconfig.jsonmoe_intermediate_size
ffns.moe.gatedtruecodemodelling codeexperts are gated MLPs
ffns.moe.shared1config.jsonconfig.jsonn_shared_experts
ffns.moe.d_shared1,280config.jsonconfig.jsonmoe_intermediate_size
layout1× full/moe · 3× kda/moe · 1× full/moe · 3× kda/moe · 1× full/moe · 3× kda/moe · 1× full/moe · 3× kda/moe · 1× full/moe · 3× kda/moe · 1× full/moe · 3× kda/moe · 1× full/moe · 3× kda/moe · 1× full/moe · 3× kda/moe · 1× full/moe · 3× kda/moe · 1× full/moe · 3× kda/moe · 1× full/moe · 3× kda/moe · 1× full/moe · 3× kda/moeconfig.jsonconfig.jsongqa_layers

Sources

Listed in the LLM Architecture Gallery checklist as “Solar Open 2 (250B)” (name only; see about).