llm-architectures-explained

/models

LFM2.5 350M

Liquid AI · LFM2 · open weights

Facts and where they come from

Released2026-03config.jsonconfig.jsonHugging Face repository creation date (api.createdAt)
Licenceotherconfig.jsonconfig.jsonREADME metadata: license
Total parameters350Mlabmodel cardREADME: Number of parameters: 350M
Active parametersnot disclosednot disclosed
Context length32K tokenslabmodel cardREADME: Context length: 32,768 tokens
Norm placementprecodemodelling codetransformers 5.18.0 lfm2: operator_norm and ffn_norm
Norm typeRMSNormcodemodelling codetransformers 5.18.0 lfm2: operator_norm and ffn_norm
QK-normyescodemodelling codetransformers 5.18.0 lfm2: q_layernorm and k_layernorm
Positional encodingRoPEcodemodelling coderotary on the full head (default)
Parallel attention and MLPnocodemodelling codetransformers 5.18.0 lfm2: operator_norm and ffn_norm

Architecture, drawn from the data

10× Short convolution + 6× GQA 16q/8kv. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

LFM2.5 350M: layer stack and blockslayers (16)mixer / FFNlayer 0: Gated short convolution · kernel 3layer 0: Gated MLP: 4608layer 1: Gated short convolution · kernel 3layer 1: Gated MLP: 4608layer 2: GQA: 16 query / 8 KV heads · head 64layer 2: Gated MLP: 4608layer 3: Gated short convolution · kernel 3layer 3: Gated MLP: 4608layer 4: Gated short convolution · kernel 3layer 4: Gated MLP: 4608layer 5: GQA: 16 query / 8 KV heads · head 64layer 5: Gated MLP: 4608layer 6: Gated short convolution · kernel 3layer 6: Gated MLP: 4608layer 7: Gated short convolution · kernel 3layer 7: Gated MLP: 4608layer 8: GQA: 16 query / 8 KV heads · head 64layer 8: Gated MLP: 4608layer 9: Gated short convolution · kernel 3layer 9: Gated MLP: 4608layer 10: GQA: 16 query / 8 KV heads · head 64layer 10: Gated MLP: 4608layer 11: Gated short convolution · kernel 3layer 11: Gated MLP: 4608layer 12: GQA: 16 query / 8 KV heads · head 64layer 12: Gated MLP: 4608layer 13: Gated short convolution · kernel 3layer 13: Gated MLP: 4608layer 14: GQA: 16 query / 8 KV heads · head 64layer 14: Gated MLP: 4608layer 15: Gated short convolution · kernel 3layer 15: Gated MLP: 46080815× 10normGated short convolution · kernel 3+normGated MLP: 4608+× 6normGQA: 16 query / 8 KV heads · head 64+normGated MLP: 4608+short convfull attentiondense FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)354M
Active per token (modelled)354M
Without embeddings and output head287M total, 287M active
Published weights (Hugging Face count)354M
KV cache per token, BF16 (layers that grow with context)12 KiB
KV cache + state at 32K tokens, BF16384 MiB
Decode FLOPs per token at 4K context810 MFLOP
Prefill FLOPs for a 4K prompt2.56 TFLOP

KV cache against context

LFM2.5 350M: KV cache bytes against context length101001,00010,00098 KiB980 KiB9.5 MiB95 MiBcontext (tokens)KV cache + state (BF16)LFM2.5 350M

Compare with other models →

Every architecture field

FieldValueSource
d_model1,024config.jsonconfig.jsonhidden_size
vocab65,536config.jsonconfig.jsonvocab_size
tied_embeddingstrueconfig.jsonconfig.jsontie_embedding
mixers.full.typeattncodemodelling codeattention block
mixers.full.heads16config.jsonconfig.jsonnum_attention_heads
mixers.full.kv_heads8config.jsonconfig.jsonnum_key_value_heads
mixers.full.head_dim64codemodelling codetransformers 5.18.0: head_dim = hidden_size / num_attention_heads
mixers.full.qk_normtruecodemodelling codetransformers 5.18.0 lfm2: q_layernorm and k_layernorm
mixers.conv.typeconvcodemodelling codegated short convolution
mixers.conv.kernel3config.jsonconfig.jsonconv_L_cache
ffns.dense.typedensecodemodelling codeMLP
ffns.dense.d_ff4,608codemodelling codetransformers 5.18.0 lfm2: block_auto_adjust_ff_dim: int(2/3 * block_ff_dim * multiplier) rounded up to block_multiple_of
ffns.dense.gatedtruecodemodelling codeSwiGLU
layout2× conv/dense · 1× full/dense · 2× conv/dense · 1× full/dense · 2× conv/dense · 1× full/dense · 1× conv/dense · 1× full/dense · 1× conv/dense · 1× full/dense · 1× conv/dense · 1× full/dense · 1× conv/denseconfig.jsonconfig.jsonlayer_types

Sources

Listed in the LLM Architecture Gallery checklist as “LFM2.5 (350M)” (name only; see about).