llm-architectures-explained

/models

Tiny Aya

Cohere · Aya · open weights

Facts and where they come from

Released2026-02config.jsonconfig.jsonHugging Face repository creation date (api.createdAt)
Licencecc-by-nc-4.0config.jsonconfig.jsonREADME metadata: license
Total parameters3.35Blabmodel cardREADME: a pretrained 3.35 billion parameter model
Active parametersnot disclosednot disclosed
Context length8K tokenspaperpaperTable 2: input context 8192
Norm placementparallelcodemodelling codetransformers 5.18.0 cohere2: a single input_layernorm feeds attention and the MLP
Norm typeLayerNormcodemodelling codetransformers 5.18.0 cohere2: a single input_layernorm feeds attention and the MLP
QK-normnocodemodelling codeno q/k normalisation in the attention block
Positional encodingRoPE; full-attention layers have no RoPE; sliding-window layers use itpaperpaper§3.1: RoPE on sliding-window layers, NoPE on full-attention layers
Parallel attention and MLPyescodemodelling codetransformers 5.18.0 cohere2: a single input_layernorm feeds attention and the MLP

Architecture, drawn from the data

27× GQA 16q/4kv, window 4096 + 9× GQA 16q/4kv. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

Tiny Aya: layer stack and blockslayers (36)mixer / FFNlayer 0: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 0: Gated MLP: 11008layer 1: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 1: Gated MLP: 11008layer 2: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 2: Gated MLP: 11008layer 3: GQA: 16 query / 4 KV heads · head 128layer 3: Gated MLP: 11008layer 4: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 4: Gated MLP: 11008layer 5: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 5: Gated MLP: 11008layer 6: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 6: Gated MLP: 11008layer 7: GQA: 16 query / 4 KV heads · head 128layer 7: Gated MLP: 11008layer 8: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 8: Gated MLP: 11008layer 9: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 9: Gated MLP: 11008layer 10: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 10: Gated MLP: 11008layer 11: GQA: 16 query / 4 KV heads · head 128layer 11: Gated MLP: 11008layer 12: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 12: Gated MLP: 11008layer 13: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 13: Gated MLP: 11008layer 14: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 14: Gated MLP: 11008layer 15: GQA: 16 query / 4 KV heads · head 128layer 15: Gated MLP: 11008layer 16: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 16: Gated MLP: 11008layer 17: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 17: Gated MLP: 11008layer 18: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 18: Gated MLP: 11008layer 19: GQA: 16 query / 4 KV heads · head 128layer 19: Gated MLP: 11008layer 20: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 20: Gated MLP: 11008layer 21: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 21: Gated MLP: 11008layer 22: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 22: Gated MLP: 11008layer 23: GQA: 16 query / 4 KV heads · head 128layer 23: Gated MLP: 11008layer 24: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 24: Gated MLP: 11008layer 25: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 25: Gated MLP: 11008layer 26: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 26: Gated MLP: 11008layer 27: GQA: 16 query / 4 KV heads · head 128layer 27: Gated MLP: 11008layer 28: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 28: Gated MLP: 11008layer 29: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 29: Gated MLP: 11008layer 30: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 30: Gated MLP: 11008layer 31: GQA: 16 query / 4 KV heads · head 128layer 31: Gated MLP: 11008layer 32: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 32: Gated MLP: 11008layer 33: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 33: Gated MLP: 11008layer 34: GQA: 16 query / 4 KV heads · head 128 · window 4,096layer 34: Gated MLP: 11008layer 35: GQA: 16 query / 4 KV heads · head 128layer 35: Gated MLP: 1100801835× 27normGQA: 16 query / 4 KV heads · head 128 · window 4,096Gated MLP: 11008+× 9normGQA: 16 query / 4 KV heads · head 128Gated MLP: 11008+sliding windowfull attentiondense FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)3.35B
Active per token (modelled)3.35B
Without embeddings and output head2.81B total, 2.81B active
Published weights (Hugging Face count)3.35B
KV cache per token, BF16 (layers that grow with context)18 KiB
KV cache + state at 8K tokens, BF16360 MiB
Decode FLOPs per token at 4K context7.91 GFLOP
Prefill FLOPs for a 4K prompt25.5 TFLOP

KV cache against context

Tiny Aya: KV cache bytes against context length101001,000980 KiB9.5 MiB95 MiBcontext (tokens)KV cache + state (BF16)Tiny Aya

Compare with other models →

Every architecture field

FieldValueSource
d_model2,048paperpaperTable 2 (p. 9): embedding dims 2048
vocab262,144paperpaperTable 2: vocab size 262k (262,144 assumed)
tied_embeddingstruepaperpaperTable 2: 0.5B embedding parameters counted once (tied)
mixers.full.typeattncodemodelling codeattention block
mixers.full.heads16paperpaperTable 2: num heads 16
mixers.full.kv_heads4paperpaperTable 2: num KV heads 4
mixers.full.head_dim128codemodelling codetransformers 5.18.0: head_dim = hidden_size / num_attention_heads
mixers.sliding.typeattncodemodelling codeattention block
mixers.sliding.heads16paperpaperTable 2: num heads 16
mixers.sliding.kv_heads4paperpaperTable 2: num KV heads 4
mixers.sliding.head_dim128codemodelling codetransformers 5.18.0: head_dim = hidden_size / num_attention_heads
mixers.sliding.window4,096paperpaperTable 2: sliding window 4096
ffns.dense.typedensecodemodelling codeMLP block
ffns.dense.d_ff11,008paperpaperTable 2: FFN hidden dims 11008
ffns.dense.gatedtruecodemodelling codeMLP: gated (SwiGLU/GeGLU)
layout3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/dense · 3× sliding/dense · 1× full/densepaperpaper§3.1: sliding window and full attention in a 3:1 ratio
norms_per_layer1codemodelling codeparallel block: one norm feeds attention and MLP

Sources

Listed in the LLM Architecture Gallery checklist as “Tiny Aya (3.35B)” (name only; see about).