llm-architectures-explained

/models

GPT-J 6B

EleutherAI · GPT-J · open weights

Facts and where they come from

Released2021-06labmodel cardREADME citation: GPT-J-6B, June 2021
Licenceapache-2.0config.jsonconfig.jsonREADME metadata: license
Total parameters6Blabmodel cardREADME: GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Active parametersnot disclosednot disclosed
Context length2K tokensconfig.jsonconfig.jsonn_positions
Norm placementparallelcodemodelling codetransformers 5.18.0 gptj: one ln_1 feeds attention and the MLP in parallel
Norm typeLayerNormcodemodelling codetransformers 5.18.0 gptj: one ln_1 feeds attention and the MLP in parallel
QK-normnocodemodelling codeno q/k normalisation in the attention block
Positional encodingRoPEcodemodelling coderotary on the full head (default)
Parallel attention and MLPyescodemodelling codetransformers 5.18.0 gptj: one ln_1 feeds attention and the MLP in parallel

Architecture, drawn from the data

MHA 16q/16kv. Each column is one layer: its token mixer above, its feed-forward block below. Paler columns reuse another layer’s keys and values.

GPT-J 6B: layer stack and blockslayers (28)mixer / FFNlayer 0: MHA: 16 query / 16 KV heads · head 256layer 0: MLP: 16384layer 1: MHA: 16 query / 16 KV heads · head 256layer 1: MLP: 16384layer 2: MHA: 16 query / 16 KV heads · head 256layer 2: MLP: 16384layer 3: MHA: 16 query / 16 KV heads · head 256layer 3: MLP: 16384layer 4: MHA: 16 query / 16 KV heads · head 256layer 4: MLP: 16384layer 5: MHA: 16 query / 16 KV heads · head 256layer 5: MLP: 16384layer 6: MHA: 16 query / 16 KV heads · head 256layer 6: MLP: 16384layer 7: MHA: 16 query / 16 KV heads · head 256layer 7: MLP: 16384layer 8: MHA: 16 query / 16 KV heads · head 256layer 8: MLP: 16384layer 9: MHA: 16 query / 16 KV heads · head 256layer 9: MLP: 16384layer 10: MHA: 16 query / 16 KV heads · head 256layer 10: MLP: 16384layer 11: MHA: 16 query / 16 KV heads · head 256layer 11: MLP: 16384layer 12: MHA: 16 query / 16 KV heads · head 256layer 12: MLP: 16384layer 13: MHA: 16 query / 16 KV heads · head 256layer 13: MLP: 16384layer 14: MHA: 16 query / 16 KV heads · head 256layer 14: MLP: 16384layer 15: MHA: 16 query / 16 KV heads · head 256layer 15: MLP: 16384layer 16: MHA: 16 query / 16 KV heads · head 256layer 16: MLP: 16384layer 17: MHA: 16 query / 16 KV heads · head 256layer 17: MLP: 16384layer 18: MHA: 16 query / 16 KV heads · head 256layer 18: MLP: 16384layer 19: MHA: 16 query / 16 KV heads · head 256layer 19: MLP: 16384layer 20: MHA: 16 query / 16 KV heads · head 256layer 20: MLP: 16384layer 21: MHA: 16 query / 16 KV heads · head 256layer 21: MLP: 16384layer 22: MHA: 16 query / 16 KV heads · head 256layer 22: MLP: 16384layer 23: MHA: 16 query / 16 KV heads · head 256layer 23: MLP: 16384layer 24: MHA: 16 query / 16 KV heads · head 256layer 24: MLP: 16384layer 25: MHA: 16 query / 16 KV heads · head 256layer 25: MLP: 16384layer 26: MHA: 16 query / 16 KV heads · head 256layer 26: MLP: 16384layer 27: MHA: 16 query / 16 KV heads · head 256layer 27: MLP: 1638401427× 28normMHA: 16 query / 16 KV heads · head 256MLP: 16384+full attentiondense FFN

Modelled costs

From the cost model, batch size 1. Totals the lab states are in the table above; differences come from rounding, from what a lab counts, or from parts the model does not describe (listed on the about page).

Parameters (modelled)6.05B
Active per token (modelled)6.05B
Without embeddings and output head5.64B total, 5.64B active
KV cache per token, BF16 (layers that grow with context)448 KiB
KV cache + state at 2K tokens, BF16896 MiB
Decode FLOPs per token at 4K context13.6 GFLOP
Prefill FLOPs for a 4K prompt50 TFLOP

KV cache against context

GPT-J 6B: KV cache bytes against context length101001,000980 KiB9.5 MiB95 MiB950 MiBcontext (tokens)KV cache + state (BF16)GPT-J 6B

Compare with other models →

Every architecture field

FieldValueSource
d_model4,096config.jsonconfig.jsonn_embd
vocab50,400config.jsonconfig.jsonvocab_size
tied_embeddingsfalseconfig.jsonconfig.jsontie_word_embeddings
mixers.full.typeattncodemodelling codeattention block
mixers.full.heads16config.jsonconfig.jsonn_head
mixers.full.kv_heads16config.jsonconfig.jsonn_head
mixers.full.head_dim256codemodelling codetransformers 5.18.0: head_dim = hidden_size / n_head
ffns.dense.typedensecodemodelling codeMLP
ffns.dense.d_ff16,384codemodelling codetransformers 5.18.0 gptj: n_inner defaults to 4 * n_embd
ffns.dense.gatedfalsecodemodelling codeGELU MLP
ffns.dense.biastruecodemodelling codeMLP biases
layout28× full/denseconfig.jsonconfig.jsonn_layer
norms_per_layer1codemodelling codeparallel attention and MLP share one LayerNorm

Sources