llm-architectures-explained

/models

Ouro 2.6B Thinking

ByteDance · Ouro · open weights

Facts and where they come from

Released2025-10config.jsonconfig.jsonHugging Face repository creation date (api.createdAt)
Licenceapache-2.0config.jsonconfig.jsonREADME metadata: license
Total parameters2.6Blabmodel cardREADME: only 2.6B parameters
Active parametersnot disclosednot disclosed
Context length64K tokensconfig.jsonconfig.jsonmax_position_embeddings
Norm placementsandwichcodemodelling coderepo modeling_ouro.py: input_layernorm(_2) and post_attention_layernorm(_2)
Norm typeRMSNormcodemodelling coderepo modeling_ouro.py: input_layernorm(_2) and post_attention_layernorm(_2)
QK-normnocodemodelling codeno q/k normalisation in the attention block
Positional encodingRoPEcodemodelling coderotary on the full head (default)
Parallel attention and MLPnocodemodelling coderepo modeling_ouro.py: input_layernorm(_2) and post_attention_layernorm(_2)

Architecture, drawn from the data

MHA 16q/16kv. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

Ouro 2.6B Thinking: layer stack and blockslayers (48)mixer / FFNlayer 0: MHA: 16 query / 16 KV heads · head 128layer 0: Gated MLP: 5632layer 1: MHA: 16 query / 16 KV heads · head 128layer 1: Gated MLP: 5632layer 2: MHA: 16 query / 16 KV heads · head 128layer 2: Gated MLP: 5632layer 3: MHA: 16 query / 16 KV heads · head 128layer 3: Gated MLP: 5632layer 4: MHA: 16 query / 16 KV heads · head 128layer 4: Gated MLP: 5632layer 5: MHA: 16 query / 16 KV heads · head 128layer 5: Gated MLP: 5632layer 6: MHA: 16 query / 16 KV heads · head 128layer 6: Gated MLP: 5632layer 7: MHA: 16 query / 16 KV heads · head 128layer 7: Gated MLP: 5632layer 8: MHA: 16 query / 16 KV heads · head 128layer 8: Gated MLP: 5632layer 9: MHA: 16 query / 16 KV heads · head 128layer 9: Gated MLP: 5632layer 10: MHA: 16 query / 16 KV heads · head 128layer 10: Gated MLP: 5632layer 11: MHA: 16 query / 16 KV heads · head 128layer 11: Gated MLP: 5632layer 12: MHA: 16 query / 16 KV heads · head 128layer 12: Gated MLP: 5632layer 13: MHA: 16 query / 16 KV heads · head 128layer 13: Gated MLP: 5632layer 14: MHA: 16 query / 16 KV heads · head 128layer 14: Gated MLP: 5632layer 15: MHA: 16 query / 16 KV heads · head 128layer 15: Gated MLP: 5632layer 16: MHA: 16 query / 16 KV heads · head 128layer 16: Gated MLP: 5632layer 17: MHA: 16 query / 16 KV heads · head 128layer 17: Gated MLP: 5632layer 18: MHA: 16 query / 16 KV heads · head 128layer 18: Gated MLP: 5632layer 19: MHA: 16 query / 16 KV heads · head 128layer 19: Gated MLP: 5632layer 20: MHA: 16 query / 16 KV heads · head 128layer 20: Gated MLP: 5632layer 21: MHA: 16 query / 16 KV heads · head 128layer 21: Gated MLP: 5632layer 22: MHA: 16 query / 16 KV heads · head 128layer 22: Gated MLP: 5632layer 23: MHA: 16 query / 16 KV heads · head 128layer 23: Gated MLP: 5632layer 24: MHA: 16 query / 16 KV heads · head 128layer 24: Gated MLP: 5632layer 25: MHA: 16 query / 16 KV heads · head 128layer 25: Gated MLP: 5632layer 26: MHA: 16 query / 16 KV heads · head 128layer 26: Gated MLP: 5632layer 27: MHA: 16 query / 16 KV heads · head 128layer 27: Gated MLP: 5632layer 28: MHA: 16 query / 16 KV heads · head 128layer 28: Gated MLP: 5632layer 29: MHA: 16 query / 16 KV heads · head 128layer 29: Gated MLP: 5632layer 30: MHA: 16 query / 16 KV heads · head 128layer 30: Gated MLP: 5632layer 31: MHA: 16 query / 16 KV heads · head 128layer 31: Gated MLP: 5632layer 32: MHA: 16 query / 16 KV heads · head 128layer 32: Gated MLP: 5632layer 33: MHA: 16 query / 16 KV heads · head 128layer 33: Gated MLP: 5632layer 34: MHA: 16 query / 16 KV heads · head 128layer 34: Gated MLP: 5632layer 35: MHA: 16 query / 16 KV heads · head 128layer 35: Gated MLP: 5632layer 36: MHA: 16 query / 16 KV heads · head 128layer 36: Gated MLP: 5632layer 37: MHA: 16 query / 16 KV heads · head 128layer 37: Gated MLP: 5632layer 38: MHA: 16 query / 16 KV heads · head 128layer 38: Gated MLP: 5632layer 39: MHA: 16 query / 16 KV heads · head 128layer 39: Gated MLP: 5632layer 40: MHA: 16 query / 16 KV heads · head 128layer 40: Gated MLP: 5632layer 41: MHA: 16 query / 16 KV heads · head 128layer 41: Gated MLP: 5632layer 42: MHA: 16 query / 16 KV heads · head 128layer 42: Gated MLP: 5632layer 43: MHA: 16 query / 16 KV heads · head 128layer 43: Gated MLP: 5632layer 44: MHA: 16 query / 16 KV heads · head 128layer 44: Gated MLP: 5632layer 45: MHA: 16 query / 16 KV heads · head 128layer 45: Gated MLP: 5632layer 46: MHA: 16 query / 16 KV heads · head 128layer 46: Gated MLP: 5632layer 47: MHA: 16 query / 16 KV heads · head 128layer 47: Gated MLP: 563202447× 48normMHA: 16 query / 16 KV heads · head 128norm+normGated MLP: 5632norm+full attentiondense FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)2.67B
Active per token (modelled)2.67B
Without embeddings and output head2.47B total, 2.47B active
Published weights (Hugging Face count)2.67B
KV cache per token, BF16 (layers that grow with context)384 KiB
KV cache + state at 64K tokens, BF1624 GiB
Decode FLOPs per token at 4K context26.4 GFLOP
Prefill FLOPs for a 4K prompt94 TFLOP

KV cache against context

Ouro 2.6B Thinking: KV cache bytes against context length101001,00010,000980 KiB9.5 MiB95 MiB950 MiB9.3 GiBcontext (tokens)KV cache + state (BF16)Ouro 2.6B Thinking

Compare with other models →

Every architecture field

FieldValueSource
d_model2,048config.jsonconfig.jsonhidden_size
vocab49,152config.jsonconfig.jsonvocab_size
tied_embeddingsfalseconfig.jsonconfig.jsontie_word_embeddings
mixers.full.typeattncodemodelling codeattention block
mixers.full.heads16config.jsonconfig.jsonnum_attention_heads
mixers.full.kv_heads16config.jsonconfig.jsonnum_key_value_heads
mixers.full.head_dim128config.jsonconfig.jsonhead_dim
mixers.full.qk_normtruecodemodelling codetransformers 5.18.0 ouro: q_norm and k_norm
ffns.dense.typedensecodemodelling codeMLP block
ffns.dense.d_ff5,632config.jsonconfig.jsonintermediate_size
ffns.dense.gatedtruecodemodelling codeMLP: gated (SwiGLU/GeGLU)
layout48× full/denseconfig.jsonconfig.jsonnum_hidden_layers
loops4config.jsonconfig.jsontotal_ut_steps

Sources

Listed in the LLM Architecture Gallery checklist as “Ouro-Thinking (2.6B)” (name only; see about).