llm-architectures-explained

/models

LFM2.5 8B-A1B

Liquid AI · LFM2 · open weights

Facts and where they come from

Released2026-05config.jsonconfig.jsonHugging Face repository creation date (api.createdAt)
Licenceotherconfig.jsonconfig.jsonREADME metadata: license
Total parameters8.3Blabmodel cardREADME: 8.3B total / 1.5B active
Active parameters1.5Blabmodel cardREADME: 1.5B active
Context length125K tokensconfig.jsonconfig.jsonmax_position_embeddings
Norm placementprecodemodelling codetransformers 5.18.0 lfm2_moe: operator_norm and ffn_norm
Norm typeRMSNormcodemodelling codetransformers 5.18.0 lfm2_moe: operator_norm and ffn_norm
QK-normyescodemodelling codetransformers 5.18.0 lfm2_moe: q_layernorm and k_layernorm
Positional encodingRoPEcodemodelling coderotary on the full head (default)
Parallel attention and MLPnocodemodelling codetransformers 5.18.0 lfm2_moe: operator_norm and ffn_norm

Architecture, drawn from the data

18× Short convolution + 6× GQA 32q/8kv. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

LFM2.5 8B-A1B: layer stack and blockslayers (24)mixer / FFNlayer 0: Gated short convolution · kernel 3layer 0: Gated MLP: 7168layer 1: Gated short convolution · kernel 3layer 1: Gated MLP: 7168layer 2: GQA: 32 query / 8 KV heads · head 64layer 2: MoE: 32 experts, 4 active · expert 1792layer 3: Gated short convolution · kernel 3layer 3: MoE: 32 experts, 4 active · expert 1792layer 4: Gated short convolution · kernel 3layer 4: MoE: 32 experts, 4 active · expert 1792layer 5: Gated short convolution · kernel 3layer 5: MoE: 32 experts, 4 active · expert 1792layer 6: GQA: 32 query / 8 KV heads · head 64layer 6: MoE: 32 experts, 4 active · expert 1792layer 7: Gated short convolution · kernel 3layer 7: MoE: 32 experts, 4 active · expert 1792layer 8: Gated short convolution · kernel 3layer 8: MoE: 32 experts, 4 active · expert 1792layer 9: Gated short convolution · kernel 3layer 9: MoE: 32 experts, 4 active · expert 1792layer 10: GQA: 32 query / 8 KV heads · head 64layer 10: MoE: 32 experts, 4 active · expert 1792layer 11: Gated short convolution · kernel 3layer 11: MoE: 32 experts, 4 active · expert 1792layer 12: Gated short convolution · kernel 3layer 12: MoE: 32 experts, 4 active · expert 1792layer 13: Gated short convolution · kernel 3layer 13: MoE: 32 experts, 4 active · expert 1792layer 14: GQA: 32 query / 8 KV heads · head 64layer 14: MoE: 32 experts, 4 active · expert 1792layer 15: Gated short convolution · kernel 3layer 15: MoE: 32 experts, 4 active · expert 1792layer 16: Gated short convolution · kernel 3layer 16: MoE: 32 experts, 4 active · expert 1792layer 17: Gated short convolution · kernel 3layer 17: MoE: 32 experts, 4 active · expert 1792layer 18: GQA: 32 query / 8 KV heads · head 64layer 18: MoE: 32 experts, 4 active · expert 1792layer 19: Gated short convolution · kernel 3layer 19: MoE: 32 experts, 4 active · expert 1792layer 20: Gated short convolution · kernel 3layer 20: MoE: 32 experts, 4 active · expert 1792layer 21: GQA: 32 query / 8 KV heads · head 64layer 21: MoE: 32 experts, 4 active · expert 1792layer 22: Gated short convolution · kernel 3layer 22: MoE: 32 experts, 4 active · expert 1792layer 23: Gated short convolution · kernel 3layer 23: MoE: 32 experts, 4 active · expert 179201223× 16normGated short convolution · kernel 3+normMoE: 32 experts, 4 active · expert 1792+× 6normGQA: 32 query / 8 KV heads · head 64+normMoE: 32 experts, 4 active · expert 1792+× 2normGated short convolution · kernel 3+normGated MLP: 7168+short convfull attentiondense FFNMoE FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)8.47B
Active per token (modelled)1.69B
Without embeddings and output head8.21B total, 1.42B active
Published weights (Hugging Face count)8.47B
KV cache per token, BF16 (layers that grow with context)12 KiB
KV cache + state at 125K tokens, BF161.46 GiB
Decode FLOPs per token at 4K context3.57 GFLOP
Prefill FLOPs for a 4K prompt12.1 TFLOP

KV cache against context

LFM2.5 8B-A1B: KV cache bytes against context length101001,00010,000980 KiB9.5 MiB95 MiBcontext (tokens)KV cache + state (BF16)LFM2.5 8B-A1B

Compare with other models →

Every architecture field

FieldValueSource
d_model2,048config.jsonconfig.jsonhidden_size
vocab128,000config.jsonconfig.jsonvocab_size
tied_embeddingstruecodemodelling codetransformers 5.18.0 lfm2_moe: tie_word_embeddings default true
mixers.full.typeattncodemodelling codeattention block
mixers.full.heads32config.jsonconfig.jsonnum_attention_heads
mixers.full.kv_heads8config.jsonconfig.jsonnum_key_value_heads
mixers.full.head_dim64codemodelling codetransformers 5.18.0: head_dim = hidden_size / num_attention_heads
mixers.full.qk_normtruecodemodelling codetransformers 5.18.0 lfm2: q_layernorm and k_layernorm
mixers.conv.typeconvcodemodelling codegated short convolution
mixers.conv.kernel3config.jsonconfig.jsonconv_L_cache
ffns.dense.typedensecodemodelling codeMLP block
ffns.dense.d_ff7,168config.jsonconfig.jsonintermediate_size
ffns.dense.gatedtruecodemodelling codeMLP: gated (SwiGLU/GeGLU)
ffns.moe.typemoecodemodelling codeMoE block
ffns.moe.experts32config.jsonconfig.jsonnum_experts
ffns.moe.active4config.jsonconfig.jsonnum_experts_per_tok
ffns.moe.d_expert1,792config.jsonconfig.jsonmoe_intermediate_size
ffns.moe.gatedtruecodemodelling codeexperts are gated MLPs
layout2× conv/dense · 1× full/moe · 3× conv/moe · 1× full/moe · 3× conv/moe · 1× full/moe · 3× conv/moe · 1× full/moe · 3× conv/moe · 1× full/moe · 2× conv/moe · 1× full/moe · 2× conv/moeconfig.jsonconfig.jsonlayer_types

Sources

Listed in the LLM Architecture Gallery checklist as “LFM2.5 (8B-A1B)” (name only; see about).