llm-architectures-explained

/models

RWKV-4 14B

RWKV · RWKV · open weights

Facts and where they come from

Released2023-05paperarXiv 2305.13048arXiv v1, May 2023
Licencenot disclosednot disclosed
Total parameters14Blabmodel cardmodel name rwkv-4-14b
Active parametersnot disclosednot disclosed
Context length1K tokensconfig.jsonconfig.jsoncontext_length
Norm placementprecodemodelling codetransformers 5.18.0 rwkv: ln1 before time mixing, ln2 before channel mixing
Norm typeLayerNormcodemodelling codetransformers 5.18.0 rwkv: ln1 before time mixing, ln2 before channel mixing
QK-normnocodemodelling codeno q/k normalisation in the attention block
Positional encodingnone; all layers: the recurrence carries ordercodemodelling codetransformers 5.18.0 rwkv: no positional encoding
Parallel attention and MLPnocodemodelling codetransformers 5.18.0 rwkv: ln1 before time mixing, ln2 before channel mixing

Architecture, drawn from the data

RWKV. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

RWKV-4 14B: layer stack and blockslayers (40)mixer / FFNlayer 0: RWKV time mixing · head 5120layer 0: MLP: 20480layer 1: RWKV time mixing · head 5120layer 1: MLP: 20480layer 2: RWKV time mixing · head 5120layer 2: MLP: 20480layer 3: RWKV time mixing · head 5120layer 3: MLP: 20480layer 4: RWKV time mixing · head 5120layer 4: MLP: 20480layer 5: RWKV time mixing · head 5120layer 5: MLP: 20480layer 6: RWKV time mixing · head 5120layer 6: MLP: 20480layer 7: RWKV time mixing · head 5120layer 7: MLP: 20480layer 8: RWKV time mixing · head 5120layer 8: MLP: 20480layer 9: RWKV time mixing · head 5120layer 9: MLP: 20480layer 10: RWKV time mixing · head 5120layer 10: MLP: 20480layer 11: RWKV time mixing · head 5120layer 11: MLP: 20480layer 12: RWKV time mixing · head 5120layer 12: MLP: 20480layer 13: RWKV time mixing · head 5120layer 13: MLP: 20480layer 14: RWKV time mixing · head 5120layer 14: MLP: 20480layer 15: RWKV time mixing · head 5120layer 15: MLP: 20480layer 16: RWKV time mixing · head 5120layer 16: MLP: 20480layer 17: RWKV time mixing · head 5120layer 17: MLP: 20480layer 18: RWKV time mixing · head 5120layer 18: MLP: 20480layer 19: RWKV time mixing · head 5120layer 19: MLP: 20480layer 20: RWKV time mixing · head 5120layer 20: MLP: 20480layer 21: RWKV time mixing · head 5120layer 21: MLP: 20480layer 22: RWKV time mixing · head 5120layer 22: MLP: 20480layer 23: RWKV time mixing · head 5120layer 23: MLP: 20480layer 24: RWKV time mixing · head 5120layer 24: MLP: 20480layer 25: RWKV time mixing · head 5120layer 25: MLP: 20480layer 26: RWKV time mixing · head 5120layer 26: MLP: 20480layer 27: RWKV time mixing · head 5120layer 27: MLP: 20480layer 28: RWKV time mixing · head 5120layer 28: MLP: 20480layer 29: RWKV time mixing · head 5120layer 29: MLP: 20480layer 30: RWKV time mixing · head 5120layer 30: MLP: 20480layer 31: RWKV time mixing · head 5120layer 31: MLP: 20480layer 32: RWKV time mixing · head 5120layer 32: MLP: 20480layer 33: RWKV time mixing · head 5120layer 33: MLP: 20480layer 34: RWKV time mixing · head 5120layer 34: MLP: 20480layer 35: RWKV time mixing · head 5120layer 35: MLP: 20480layer 36: RWKV time mixing · head 5120layer 36: MLP: 20480layer 37: RWKV time mixing · head 5120layer 37: MLP: 20480layer 38: RWKV time mixing · head 5120layer 38: MLP: 20480layer 39: RWKV time mixing · head 5120layer 39: MLP: 2048002039× 40normRWKV time mixing · head 5120+normMLP: 20480+RWKVdense FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)14.1B
Active per token (modelled)14.1B
Without embeddings and output head13.6B total, 13.6B active
KV cache per token, BF16 (layers that grow with context)0 B
KV cache + state at 1K tokens, BF161.95 GiB
Decode FLOPs per token at 4K context32 GFLOP
Prefill FLOPs for a 4K prompt129 TFLOP

KV cache against context

RWKV-4 14B: KV cache bytes against context length101001,000950 MiBcontext (tokens)KV cache + state (BF16)RWKV-4 14B

Compare with other models →

Every architecture field

FieldValueSource
d_model5,120config.jsonconfig.jsonhidden_size
vocab50,277config.jsonconfig.jsonvocab_size
tied_embeddingsfalsecodemodelling codeRWKV: separate head
mixers.rwkv.typerwkvcodemodelling codeRWKV time mixing
mixers.rwkv.head_dim5,120codemodelling codeRWKV-4: a single head of width hidden_size
mixers.rwkv.mats4codemodelling codetime mixing: receptance, key, value, output
ffns.dense.typedensecodemodelling codechannel mixing
ffns.dense.d_ff20,480config.jsonconfig.jsonintermediate_size
ffns.dense.gatedfalsecodemodelling codechannel mixing: key and value matrices
ffns.dense.receptancetruecodemodelling codechannel mixing also has a d x d receptance matrix
layout40× rwkv/denseconfig.jsonconfig.jsonnum_hidden_layers

Sources