llm-architectures-explained

/models

Antares 1B

Cisco · Granite · open weights

Facts and where they come from

Released2026-07config.jsonconfig.jsonHugging Face repository creation date (api.createdAt)
Licenceapache-2.0config.jsonconfig.jsonREADME metadata: license
Total parametersnot disclosednot disclosed
Active parametersnot disclosednot disclosed
Context length128K tokenslabmodel cardmodel card: 128K context window
Norm placementprecodemodelling codetransformers 5.18.0 / repo modelling code for llama: input_layernorm before attention, post_attention_layernorm before the MLP
Norm typeRMSNormcodemodelling codetransformers 5.18.0 / repo modelling code for llama: input_layernorm before attention, post_attention_layernorm before the MLP
QK-normnocodemodelling codeno q/k normalisation in the attention block
Positional encodingRoPEcodemodelling coderotary on the full head (default)
Parallel attention and MLPnocodemodelling codetransformers 5.18.0 / repo modelling code for llama: input_layernorm before attention, post_attention_layernorm before the MLP

Architecture, drawn from the data

GQA 16q/4kv. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

Antares 1B: layer stack and blockslayers (40)mixer / FFNlayer 0: GQA: 16 query / 4 KV heads · head 128layer 0: Gated MLP: 4096layer 1: GQA: 16 query / 4 KV heads · head 128layer 1: Gated MLP: 4096layer 2: GQA: 16 query / 4 KV heads · head 128layer 2: Gated MLP: 4096layer 3: GQA: 16 query / 4 KV heads · head 128layer 3: Gated MLP: 4096layer 4: GQA: 16 query / 4 KV heads · head 128layer 4: Gated MLP: 4096layer 5: GQA: 16 query / 4 KV heads · head 128layer 5: Gated MLP: 4096layer 6: GQA: 16 query / 4 KV heads · head 128layer 6: Gated MLP: 4096layer 7: GQA: 16 query / 4 KV heads · head 128layer 7: Gated MLP: 4096layer 8: GQA: 16 query / 4 KV heads · head 128layer 8: Gated MLP: 4096layer 9: GQA: 16 query / 4 KV heads · head 128layer 9: Gated MLP: 4096layer 10: GQA: 16 query / 4 KV heads · head 128layer 10: Gated MLP: 4096layer 11: GQA: 16 query / 4 KV heads · head 128layer 11: Gated MLP: 4096layer 12: GQA: 16 query / 4 KV heads · head 128layer 12: Gated MLP: 4096layer 13: GQA: 16 query / 4 KV heads · head 128layer 13: Gated MLP: 4096layer 14: GQA: 16 query / 4 KV heads · head 128layer 14: Gated MLP: 4096layer 15: GQA: 16 query / 4 KV heads · head 128layer 15: Gated MLP: 4096layer 16: GQA: 16 query / 4 KV heads · head 128layer 16: Gated MLP: 4096layer 17: GQA: 16 query / 4 KV heads · head 128layer 17: Gated MLP: 4096layer 18: GQA: 16 query / 4 KV heads · head 128layer 18: Gated MLP: 4096layer 19: GQA: 16 query / 4 KV heads · head 128layer 19: Gated MLP: 4096layer 20: GQA: 16 query / 4 KV heads · head 128layer 20: Gated MLP: 4096layer 21: GQA: 16 query / 4 KV heads · head 128layer 21: Gated MLP: 4096layer 22: GQA: 16 query / 4 KV heads · head 128layer 22: Gated MLP: 4096layer 23: GQA: 16 query / 4 KV heads · head 128layer 23: Gated MLP: 4096layer 24: GQA: 16 query / 4 KV heads · head 128layer 24: Gated MLP: 4096layer 25: GQA: 16 query / 4 KV heads · head 128layer 25: Gated MLP: 4096layer 26: GQA: 16 query / 4 KV heads · head 128layer 26: Gated MLP: 4096layer 27: GQA: 16 query / 4 KV heads · head 128layer 27: Gated MLP: 4096layer 28: GQA: 16 query / 4 KV heads · head 128layer 28: Gated MLP: 4096layer 29: GQA: 16 query / 4 KV heads · head 128layer 29: Gated MLP: 4096layer 30: GQA: 16 query / 4 KV heads · head 128layer 30: Gated MLP: 4096layer 31: GQA: 16 query / 4 KV heads · head 128layer 31: Gated MLP: 4096layer 32: GQA: 16 query / 4 KV heads · head 128layer 32: Gated MLP: 4096layer 33: GQA: 16 query / 4 KV heads · head 128layer 33: Gated MLP: 4096layer 34: GQA: 16 query / 4 KV heads · head 128layer 34: Gated MLP: 4096layer 35: GQA: 16 query / 4 KV heads · head 128layer 35: Gated MLP: 4096layer 36: GQA: 16 query / 4 KV heads · head 128layer 36: Gated MLP: 4096layer 37: GQA: 16 query / 4 KV heads · head 128layer 37: Gated MLP: 4096layer 38: GQA: 16 query / 4 KV heads · head 128layer 38: Gated MLP: 4096layer 39: GQA: 16 query / 4 KV heads · head 128layer 39: Gated MLP: 409602039× 40normGQA: 16 query / 4 KV heads · head 128+normGated MLP: 4096+full attentiondense FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)1.84B
Active per token (modelled)1.84B
Without embeddings and output head1.43B total, 1.43B active
Published weights (Hugging Face count)1.84B
KV cache per token, BF16 (layers that grow with context)80 KiB
KV cache + state at 128K tokens, BF1610 GiB
Decode FLOPs per token at 4K context4.61 GFLOP
Prefill FLOPs for a 4K prompt14.4 TFLOP

KV cache against context

Antares 1B: KV cache bytes against context length101001,00010,000100,000980 KiB9.5 MiB95 MiB950 MiB9.3 GiBcontext (tokens)KV cache + state (BF16)Antares 1B

Compare with other models →

Every architecture field

FieldValueSource
d_model2,048labmodel cardmodel card: hidden dim 2048
vocab100,352labmodel cardmodel card: 100,352 vocab
tied_embeddingsfalselabmodel cardpublished weights (fdtn-ai/antares-1b@10417eb safetensors, 1,837,271,040 parameters) hold a separate output head; the Granite 4.0 1B backbone ties it
mixers.full.typeattncodemodelling codeattention block
mixers.full.heads16labmodel cardmodel card: 16 attention heads
mixers.full.kv_heads4labmodel cardmodel card: 4 KV heads (GQA)
mixers.full.head_dim128codemodelling codetransformers 5.18.0: head_dim = hidden_size / num_attention_heads
ffns.dense.typedensecodemodelling codeMLP block
ffns.dense.d_ff4,096labmodel cardGranite 4.0 1B backbone config (ibm-granite/granite-4.0-1b@6a7381b): shared_intermediate_size 4096
ffns.dense.gatedtruecodemodelling codeMLP: gated (SwiGLU/GeGLU)
layout40× full/denselabmodel cardmodel card: 40 layers

Sources

Listed in the LLM Architecture Gallery checklist as “Antares (1B)” (name only; see about).