llm-architectures-explained

/models

xLSTM 7B

NXAI · xLSTM · open weights

Facts and where they come from

Released2024-12config.jsonconfig.jsonHugging Face repository creation date (api.createdAt)
Licenceotherconfig.jsonconfig.jsonREADME metadata: license
Total parameters7Blabmodel cardmodel name xLSTM-7b
Active parametersnot disclosednot disclosed
Context lengthnot disclosednot disclosed
Norm placementprecodemodelling codetransformers 5.18.0 xlstm: norm_mlstm and norm_ffn before each block
Norm typeRMSNormcodemodelling codetransformers 5.18.0 xlstm: norm_mlstm and norm_ffn before each block
QK-normnoconfig.jsonconfig.jsonadd_qk_norm
Positional encodingnone; all layers: the recurrence carries ordercodemodelling codetransformers 5.18.0 xlstm: no positional encoding
Parallel attention and MLPnocodemodelling codetransformers 5.18.0 xlstm: norm_mlstm and norm_ffn before each block

Architecture, drawn from the data

mLSTM. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

xLSTM 7B: layer stack and blockslayers (32)mixer / FFNlayer 0: mLSTM: 8 headslayer 0: Gated MLP: 10944layer 1: mLSTM: 8 headslayer 1: Gated MLP: 10944layer 2: mLSTM: 8 headslayer 2: Gated MLP: 10944layer 3: mLSTM: 8 headslayer 3: Gated MLP: 10944layer 4: mLSTM: 8 headslayer 4: Gated MLP: 10944layer 5: mLSTM: 8 headslayer 5: Gated MLP: 10944layer 6: mLSTM: 8 headslayer 6: Gated MLP: 10944layer 7: mLSTM: 8 headslayer 7: Gated MLP: 10944layer 8: mLSTM: 8 headslayer 8: Gated MLP: 10944layer 9: mLSTM: 8 headslayer 9: Gated MLP: 10944layer 10: mLSTM: 8 headslayer 10: Gated MLP: 10944layer 11: mLSTM: 8 headslayer 11: Gated MLP: 10944layer 12: mLSTM: 8 headslayer 12: Gated MLP: 10944layer 13: mLSTM: 8 headslayer 13: Gated MLP: 10944layer 14: mLSTM: 8 headslayer 14: Gated MLP: 10944layer 15: mLSTM: 8 headslayer 15: Gated MLP: 10944layer 16: mLSTM: 8 headslayer 16: Gated MLP: 10944layer 17: mLSTM: 8 headslayer 17: Gated MLP: 10944layer 18: mLSTM: 8 headslayer 18: Gated MLP: 10944layer 19: mLSTM: 8 headslayer 19: Gated MLP: 10944layer 20: mLSTM: 8 headslayer 20: Gated MLP: 10944layer 21: mLSTM: 8 headslayer 21: Gated MLP: 10944layer 22: mLSTM: 8 headslayer 22: Gated MLP: 10944layer 23: mLSTM: 8 headslayer 23: Gated MLP: 10944layer 24: mLSTM: 8 headslayer 24: Gated MLP: 10944layer 25: mLSTM: 8 headslayer 25: Gated MLP: 10944layer 26: mLSTM: 8 headslayer 26: Gated MLP: 10944layer 27: mLSTM: 8 headslayer 27: Gated MLP: 10944layer 28: mLSTM: 8 headslayer 28: Gated MLP: 10944layer 29: mLSTM: 8 headslayer 29: Gated MLP: 10944layer 30: mLSTM: 8 headslayer 30: Gated MLP: 10944layer 31: mLSTM: 8 headslayer 31: Gated MLP: 1094401631× 32normmLSTM: 8 heads+normGated MLP: 10944+mLSTMdense FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)6.87B
Active per token (modelled)6.87B
Without embeddings and output head6.45B total, 6.45B active
Published weights (Hugging Face count)6.87B
KV cache per token, BF16 (layers that grow with context)0 B
KV cache + state at 128K tokens, BF1664.1 MiB
Decode FLOPs per token at 4K context13.5 GFLOP
Prefill FLOPs for a 4K prompt53.4 TFLOP

KV cache against context

xLSTM 7B: KV cache bytes against context length101001,00010,000100,00095 MiBcontext (tokens)KV cache + state (BF16)xLSTM 7B

Compare with other models →

Every architecture field

FieldValueSource
d_model4,096config.jsonconfig.jsonembedding_dim
vocab50,304config.jsonconfig.jsonvocab_size
tied_embeddingsfalsecodemodelling codetransformers 5.18.0 xlstm: separate lm_head
mixers.mlstm.typemlstmcodemodelling codemLSTM cell
mixers.mlstm.heads8config.jsonconfig.jsonnum_heads
mixers.mlstm.qk_dim2,048codemodelling codeqk_dim = round_up(embedding_dim * qk_dim_factor, mlstm_round_up_to_multiple_of)
mixers.mlstm.v_dim4,096codemodelling codev_dim = round_up(embedding_dim * v_dim_factor, mlstm_round_up_to_multiple_of)
ffns.dense.typedensecodemodelling codeMLP
ffns.dense.d_ff10,944codemodelling coded_ff = round_up(embedding_dim * ffn_proj_factor, ffn_round_up_to_multiple_of)
ffns.dense.gatedtruecodemodelling codeSwiGLU
layout32× mlstm/denseconfig.jsonconfig.jsonnum_blocks

Sources

Listed in the LLM Architecture Gallery checklist as “xLSTM (7B)” (name only; see about).