Mamba 2.8B
Carnegie Mellon / Princeton · Mamba · open weights
- Mamba (SSM)
- Pre-norm
Facts and where they come from
| Released | 2023-12 | paperarXiv 2312.00752arXiv v1, December 2023 |
|---|---|---|
| Licence | not disclosed | not disclosed |
| Total parameters | 2.8B | labmodel cardmodel name mamba-2.8b |
| Active parameters | not disclosed | not disclosed |
| Context length | not disclosed | not disclosed |
| Norm placement | pre | codemodelling codetransformers 5.18.0 mamba: one norm before each block |
| Norm type | RMSNorm | codemodelling codetransformers 5.18.0 mamba: one norm before each block |
| QK-norm | no | codemodelling codeno q/k normalisation in the attention block |
| Positional encoding | none; all layers: the recurrence carries order | codemodelling codetransformers 5.18.0 mamba: no positional encoding |
| Parallel attention and MLP | no | codemodelling codetransformers 5.18.0 mamba: one norm before each block |
Architecture, drawn from the data
Mamba. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.
Modelled costs
From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).
| Parameters (modelled) | 2.77B |
|---|---|
| Active per token (modelled) | 2.77B |
| Without embeddings and output head | 2.64B total, 2.64B active |
| Published weights (Hugging Face count) | 2.77B |
| KV cache per token, BF16 (layers that grow with context) | 0 B |
| KV cache + state at 128K tokens, BF16 | 11.9 MiB |
| Decode FLOPs per token at 4K context | 5.57 GFLOP |
| Prefill FLOPs for a 4K prompt | 21.7 TFLOP |
KV cache against context
Every architecture field
| Field | Value | Source |
|---|---|---|
| d_model | 2,560 | config.jsonconfig.jsonhidden_size |
| vocab | 50,280 | config.jsonconfig.jsonvocab_size |
| tied_embeddings | true | codemodelling codetransformers 5.18.0 mamba: tied |
| mixers.mamba.type | mamba1 | codemodelling codeMamba (selective SSM) |
| mixers.mamba.d_inner | 5,120 | config.jsonconfig.jsonintermediate_size |
| mixers.mamba.state | 16 | config.jsonconfig.jsonstate_size |
| mixers.mamba.conv_kernel | 4 | config.jsonconfig.jsonconv_kernel |
| mixers.mamba.dt_rank | 160 | config.jsonconfig.jsontime_step_rank |
| layout | 64× mamba/none | config.jsonconfig.jsonnum_hidden_layers |
| norms_per_layer | 1 | codemodelling codeone RMSNorm per Mamba block |
Sources
- config.json @ 96c48e0
- model card
- arXiv 2312.00752
- modelling code · transformers 5.18.0 modelling code, or the model repository's own modelling file at the pinned revision