llm-architectures-explained

/models

Step 3.5 Flash

StepFun · Step · open weights

Facts and where they come from

Released2026-02config.jsonconfig.jsonHugging Face repository creation date (api.createdAt)
Licenceapache-2.0config.jsonconfig.jsonREADME metadata: license
Total parameters196Blabmodel cardREADME: activates only 11B of its 196B parameters per token
Active parameters11Blabmodel cardREADME: 11B
Context length256K tokenslabmodel cardREADME: 256K context window
Norm placementprecodemodelling codetransformers 5.18.0 / repo modelling code for step3p5: input_layernorm before attention, post_attention_layernorm before the MLP
Norm typeRMSNormcodemodelling codetransformers 5.18.0 / repo modelling code for step3p5: input_layernorm before attention, post_attention_layernorm before the MLP
QK-normyesconfig.jsonconfig.jsonuse_qk_norm
Positional encodingRoPEcodemodelling coderotary on the full head (default)
Parallel attention and MLPnocodemodelling codetransformers 5.18.0 / repo modelling code for step3p5: input_layernorm before attention, post_attention_layernorm before the MLP

Architecture, drawn from the data

12× GQA 64q/8kv + 36× GQA 96q/8kv, window 512. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

Step 3.5 Flash: layer stack and blockslayers (48)mixer / FFNlayer 0: GQA: 64 query / 8 KV heads · head 128 · headwise output gatelayer 0: Gated MLP: 11264layer 1: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 1: Gated MLP: 11264layer 2: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 2: Gated MLP: 11264layer 3: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 3: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 4: GQA: 64 query / 8 KV heads · head 128 · headwise output gatelayer 4: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 5: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 5: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 6: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 6: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 7: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 7: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 8: GQA: 64 query / 8 KV heads · head 128 · headwise output gatelayer 8: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 9: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 9: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 10: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 10: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 11: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 11: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 12: GQA: 64 query / 8 KV heads · head 128 · headwise output gatelayer 12: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 13: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 13: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 14: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 14: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 15: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 15: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 16: GQA: 64 query / 8 KV heads · head 128 · headwise output gatelayer 16: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 17: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 17: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 18: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 18: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 19: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 19: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 20: GQA: 64 query / 8 KV heads · head 128 · headwise output gatelayer 20: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 21: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 21: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 22: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 22: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 23: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 23: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 24: GQA: 64 query / 8 KV heads · head 128 · headwise output gatelayer 24: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 25: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 25: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 26: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 26: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 27: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 27: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 28: GQA: 64 query / 8 KV heads · head 128 · headwise output gatelayer 28: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 29: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 29: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 30: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 30: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 31: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 31: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 32: GQA: 64 query / 8 KV heads · head 128 · headwise output gatelayer 32: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 33: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 33: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 34: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 34: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 35: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 35: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 36: GQA: 64 query / 8 KV heads · head 128 · headwise output gatelayer 36: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 37: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 37: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 38: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 38: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 39: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 39: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 40: GQA: 64 query / 8 KV heads · head 128 · headwise output gatelayer 40: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 41: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 41: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 42: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 42: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 43: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 43: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 44: GQA: 64 query / 8 KV heads · head 128 · headwise output gatelayer 44: MoE: 288 experts, 8 active · expert 1280 · 1 sharedlayer 45: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 45: Gated MLP: 11264layer 46: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 46: Gated MLP: 11264layer 47: GQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gatelayer 47: Gated MLP: 1126402447× 31normGQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gate+normMoE: 288 experts, 8 active · expert 1280 · 1 shared+× 11normGQA: 64 query / 8 KV heads · head 128 · headwise output gate+normMoE: 288 experts, 8 active · expert 1280 · 1 shared+× 5normGQA: 96 query / 8 KV heads · head 128 · window 512 · headwise output gate+normGated MLP: 11264+× 1normGQA: 64 query / 8 KV heads · head 128 · headwise output gate+normGated MLP: 11264+full attentionsliding windowdense FFNMoE FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)198B
Active per token (modelled)12.7B
Without embeddings and output head197B total, 11.7B active
Multi-token-prediction layers (extra)844M
Published weights (Hugging Face count)199B
KV cache per token, BF16 (layers that grow with context)48 KiB
KV cache + state at 256K tokens, BF1612.1 GiB
Decode FLOPs per token at 4K context26.9 GFLOP
Prefill FLOPs for a 4K prompt102 TFLOP

KV cache against context

Step 3.5 Flash: KV cache bytes against context length101001,00010,000100,000980 KiB9.5 MiB95 MiB950 MiB9.3 GiBcontext (tokens)KV cache + state (BF16)Step 3.5 Flash

Compare with other models →

Every architecture field

FieldValueSource
d_model4,096config.jsonconfig.jsonhidden_size
vocab128,896config.jsonconfig.jsonvocab_size
tied_embeddingsfalseconfig.jsonconfig.jsontie_word_embeddings
mixers.full.typeattncodemodelling codeattention block
mixers.full.heads64config.jsonconfig.jsonnum_attention_heads
mixers.full.kv_heads8config.jsonconfig.jsonnum_attention_groups
mixers.full.head_dim128config.jsonconfig.jsonhead_dim
mixers.full.gateheadwisecodemodelling codeuse_head_wise_attn_gate: true
mixers.full.qk_normtrueconfig.jsonconfig.jsonuse_qk_norm
mixers.sliding.typeattncodemodelling codeattention
mixers.sliding.heads96config.jsonconfig.jsonattention_other_setting.num_attention_heads
mixers.sliding.kv_heads8config.jsonconfig.jsonattention_other_setting.num_attention_groups
mixers.sliding.head_dim128config.jsonconfig.jsonattention_other_setting.head_dim
mixers.sliding.window512config.jsonconfig.jsonsliding_window
mixers.sliding.gateheadwisecodemodelling codeuse_head_wise_attn_gate: true
mixers.sliding.qk_normtrueconfig.jsonconfig.jsonuse_qk_norm
ffns.dense.typedensecodemodelling codeMLP block
ffns.dense.d_ff11,264config.jsonconfig.jsonintermediate_size
ffns.dense.gatedtruecodemodelling codeMLP: gated (SwiGLU/GeGLU)
ffns.moe.typemoecodemodelling codeMoE block
ffns.moe.experts288config.jsonconfig.jsonmoe_num_experts
ffns.moe.active8config.jsonconfig.jsonmoe_top_k
ffns.moe.d_expert1,280config.jsonconfig.jsonmoe_intermediate_size
ffns.moe.gatedtruecodemodelling codeexperts are gated MLPs
ffns.moe.shared1codemodelling codeshare_expert_dim > 0: one shared expert
ffns.moe.d_shared1,280config.jsonconfig.jsonshare_expert_dim
layout1× full/dense · 2× sliding/dense · 1× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/denseconfig.jsonconfig.jsonlayer_types, moe_layers_enum
mtp_layers3config.jsonconfig.jsonnum_nextn_predict_layers
mtp_layer{"mixer":"sliding","ffn":"dense","n":1}codemodelling codeMTP block: sliding-window attention + dense MLP

Sources

Listed in the LLM Architecture Gallery checklist as “Step 3.5 Flash (196B)” (name only; see about).