Step 3.5 Flash
StepFun · Step · open weights
- GQA
- Sliding window
- RoPE
- Pre-norm
- QK-norm
- MoE
- Shared expert
- Dense first layers
- MTP
Facts and where they come from
| Released | 2026-02 | config.jsonconfig.jsonHugging Face repository creation date (api.createdAt) |
|---|---|---|
| Licence | apache-2.0 | config.jsonconfig.jsonREADME metadata: license |
| Total parameters | 196B | labmodel cardREADME: activates only 11B of its 196B parameters per token |
| Active parameters | 11B | labmodel cardREADME: 11B |
| Context length | 256K tokens | labmodel cardREADME: 256K context window |
| Norm placement | pre | codemodelling codetransformers 5.18.0 / repo modelling code for step3p5: input_layernorm before attention, post_attention_layernorm before the MLP |
| Norm type | RMSNorm | codemodelling codetransformers 5.18.0 / repo modelling code for step3p5: input_layernorm before attention, post_attention_layernorm before the MLP |
| QK-norm | yes | config.jsonconfig.jsonuse_qk_norm |
| Positional encoding | RoPE | codemodelling coderotary on the full head (default) |
| Parallel attention and MLP | no | codemodelling codetransformers 5.18.0 / repo modelling code for step3p5: input_layernorm before attention, post_attention_layernorm before the MLP |
Architecture, drawn from the data
12× GQA 64q/8kv + 36× GQA 96q/8kv, window 512. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.
Modelled costs
From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).
| Parameters (modelled) | 198B |
|---|---|
| Active per token (modelled) | 12.7B |
| Without embeddings and output head | 197B total, 11.7B active |
| Multi-token-prediction layers (extra) | 844M |
| Published weights (Hugging Face count) | 199B |
| KV cache per token, BF16 (layers that grow with context) | 48 KiB |
| KV cache + state at 256K tokens, BF16 | 12.1 GiB |
| Decode FLOPs per token at 4K context | 26.9 GFLOP |
| Prefill FLOPs for a 4K prompt | 102 TFLOP |
KV cache against context
Every architecture field
| Field | Value | Source |
|---|---|---|
| d_model | 4,096 | config.jsonconfig.jsonhidden_size |
| vocab | 128,896 | config.jsonconfig.jsonvocab_size |
| tied_embeddings | false | config.jsonconfig.jsontie_word_embeddings |
| mixers.full.type | attn | codemodelling codeattention block |
| mixers.full.heads | 64 | config.jsonconfig.jsonnum_attention_heads |
| mixers.full.kv_heads | 8 | config.jsonconfig.jsonnum_attention_groups |
| mixers.full.head_dim | 128 | config.jsonconfig.jsonhead_dim |
| mixers.full.gate | headwise | codemodelling codeuse_head_wise_attn_gate: true |
| mixers.full.qk_norm | true | config.jsonconfig.jsonuse_qk_norm |
| mixers.sliding.type | attn | codemodelling codeattention |
| mixers.sliding.heads | 96 | config.jsonconfig.jsonattention_other_setting.num_attention_heads |
| mixers.sliding.kv_heads | 8 | config.jsonconfig.jsonattention_other_setting.num_attention_groups |
| mixers.sliding.head_dim | 128 | config.jsonconfig.jsonattention_other_setting.head_dim |
| mixers.sliding.window | 512 | config.jsonconfig.jsonsliding_window |
| mixers.sliding.gate | headwise | codemodelling codeuse_head_wise_attn_gate: true |
| mixers.sliding.qk_norm | true | config.jsonconfig.jsonuse_qk_norm |
| ffns.dense.type | dense | codemodelling codeMLP block |
| ffns.dense.d_ff | 11,264 | config.jsonconfig.jsonintermediate_size |
| ffns.dense.gated | true | codemodelling codeMLP: gated (SwiGLU/GeGLU) |
| ffns.moe.type | moe | codemodelling codeMoE block |
| ffns.moe.experts | 288 | config.jsonconfig.jsonmoe_num_experts |
| ffns.moe.active | 8 | config.jsonconfig.jsonmoe_top_k |
| ffns.moe.d_expert | 1,280 | config.jsonconfig.jsonmoe_intermediate_size |
| ffns.moe.gated | true | codemodelling codeexperts are gated MLPs |
| ffns.moe.shared | 1 | codemodelling codeshare_expert_dim > 0: one shared expert |
| ffns.moe.d_shared | 1,280 | config.jsonconfig.jsonshare_expert_dim |
| layout | 1× full/dense · 2× sliding/dense · 1× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/moe · 1× full/moe · 3× sliding/dense | config.jsonconfig.jsonlayer_types, moe_layers_enum |
| mtp_layers | 3 | config.jsonconfig.jsonnum_nextn_predict_layers |
| mtp_layer | {"mixer":"sliding","ffn":"dense","n":1} | codemodelling codeMTP block: sliding-window attention + dense MLP |
Sources
- config.json @ ab446a3
- model card
- arXiv 2602.10604
- modelling code · transformers 5.18.0 modelling code, or the model repository's own modelling file at the pinned revision
Listed in the LLM Architecture Gallery checklist as “Step 3.5 Flash (196B)” (name only; see about).