llm-architectures-explained

/models

Mistral 7B

Mistral AI · Mistral · open weights

Facts and where they come from

Released2023-09labmodel cardHF repository created September 2023
Licenceapache-2.0config.jsonconfig.jsonREADME metadata: license
Total parameters7Blabmodel cardmodel name Mistral-7B
Active parametersnot disclosednot disclosed
Context length32K tokensconfig.jsonconfig.jsonmax_position_embeddings
Norm placementprecodemodelling codetransformers 5.18.0 / repo modelling code for mistral: input_layernorm before attention, post_attention_layernorm before the MLP
Norm typeRMSNormcodemodelling codetransformers 5.18.0 / repo modelling code for mistral: input_layernorm before attention, post_attention_layernorm before the MLP
QK-normnocodemodelling codeno q/k normalisation in the attention block
Positional encodingRoPEcodemodelling coderotary on the full head (default)
Parallel attention and MLPnocodemodelling codetransformers 5.18.0 / repo modelling code for mistral: input_layernorm before attention, post_attention_layernorm before the MLP

Architecture, drawn from the data

GQA 32q/8kv, window 4096. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

Mistral 7B: layer stack and blockslayers (32)mixer / FFNlayer 0: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 0: Gated MLP: 14336layer 1: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 1: Gated MLP: 14336layer 2: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 2: Gated MLP: 14336layer 3: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 3: Gated MLP: 14336layer 4: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 4: Gated MLP: 14336layer 5: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 5: Gated MLP: 14336layer 6: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 6: Gated MLP: 14336layer 7: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 7: Gated MLP: 14336layer 8: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 8: Gated MLP: 14336layer 9: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 9: Gated MLP: 14336layer 10: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 10: Gated MLP: 14336layer 11: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 11: Gated MLP: 14336layer 12: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 12: Gated MLP: 14336layer 13: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 13: Gated MLP: 14336layer 14: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 14: Gated MLP: 14336layer 15: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 15: Gated MLP: 14336layer 16: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 16: Gated MLP: 14336layer 17: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 17: Gated MLP: 14336layer 18: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 18: Gated MLP: 14336layer 19: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 19: Gated MLP: 14336layer 20: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 20: Gated MLP: 14336layer 21: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 21: Gated MLP: 14336layer 22: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 22: Gated MLP: 14336layer 23: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 23: Gated MLP: 14336layer 24: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 24: Gated MLP: 14336layer 25: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 25: Gated MLP: 14336layer 26: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 26: Gated MLP: 14336layer 27: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 27: Gated MLP: 14336layer 28: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 28: Gated MLP: 14336layer 29: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 29: Gated MLP: 14336layer 30: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 30: Gated MLP: 14336layer 31: GQA: 32 query / 8 KV heads · head 128 · window 4,096layer 31: Gated MLP: 1433601631× 32normGQA: 32 query / 8 KV heads · head 128 · window 4,096+normGated MLP: 14336+sliding windowdense FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)7.24B
Active per token (modelled)7.24B
Without embeddings and output head6.98B total, 6.98B active
Published weights (Hugging Face count)7.24B
KV cache per token, BF16 (layers that grow with context)0 B
KV cache + state at 32K tokens, BF16512 MiB
Decode FLOPs per token at 4K context16.4 GFLOP
Prefill FLOPs for a 4K prompt61.6 TFLOP

KV cache against context

Mistral 7B: KV cache bytes against context length101001,00010,000980 KiB9.5 MiB95 MiBcontext (tokens)KV cache + state (BF16)Mistral 7B

Compare with other models →

Every architecture field

FieldValueSource
d_model4,096config.jsonconfig.jsonhidden_size
vocab32,000config.jsonconfig.jsonvocab_size
tied_embeddingsfalseconfig.jsonconfig.jsontie_word_embeddings
mixers.full.typeattncodemodelling codeattention block
mixers.full.heads32config.jsonconfig.jsonnum_attention_heads
mixers.full.kv_heads8config.jsonconfig.jsonnum_key_value_heads
mixers.full.head_dim128codemodelling codetransformers 5.18.0: head_dim = hidden_size / num_attention_heads
mixers.full.window4,096config.jsonconfig.jsonsliding_window
ffns.dense.typedensecodemodelling codeMLP block
ffns.dense.d_ff14,336config.jsonconfig.jsonintermediate_size
ffns.dense.gatedtruecodemodelling codeMLP: gated (SwiGLU/GeGLU)
layout32× full/denseconfig.jsonconfig.jsonnum_hidden_layers

Sources