Concept
A language model is trained to predict the next token. Multi-token prediction (MTP) trains it to predict the next few as well. Gloeckle et al. (2024) did it with independent output heads on a shared trunk, as an auxiliary training task, and measured better sample efficiency. DeepSeek-V3 made the extra predictions sequential: its MTP module is one more Transformer block that takes the main model's state at depth and the embedding of the next token, and predicts one token further, so each prediction keeps the full causal chain (section 2.2, p. 10). The modules share the main model's embedding and output head.
The architecture choice shows up twice:
- In training, as a denser signal. DeepSeek-V3's report says the MTP modules can simply be discarded at inference, leaving the main model unchanged (p. 11).
- At inference, as a built-in draft model for speculative decoding. The modules propose the next tokens, the main model checks them in its next step, and every accepted draft is a token for free. DeepSeek-V3 reports that its second-token prediction is accepted 85% to 90% of the time, giving 1.8 times the tokens per second (section 5.4.3, p. 35).
The interactive puts numbers on both sides for any model in the data set that ships MTP layers. DeepSeek-V3's single module touches 688M weights per token, under 2% of the 37.6B the main model activates (both counted by this site's cost model). If each drafted token is accepted with probability , a step emits tokens with one module (1.85 at ). With decode memory-bound at batch 1, the step costs about what the main model's step costs plus the module's weights, so the idealised speed-up is 1.82× at and 1.87× at , the same neighbourhood as the 1.8× DeepSeek reports.
Loading the interactive…
Concept
Things to try. Raise the depth: with , the third and fourth drafted tokens are worth less and less (, ), while every module adds its weights to every step. Then pick a model with several MTP layers (Step 3.5 Flash ships three) and compare.
MTP: 37 models in the data set
Multi-token prediction layers. Newest first; each links to its sourced page.
- MiMo-V2.6-Flash-RL
- DeepSeek-V4.1-Flash
- Qwen3.8-Flash-Next
- Qwen3.8 27B
- Ornith 1.5 35B-A3B
- Nemotron 3.5 Lightning
- Ling 3.0 Flash
- Hy4 preview
- GLM-5.3-Flash
- Motif 3 Beta
- Inkling
- BTL-3
- Nemotron 3 Ultra
- MiniMax-M3
- GLM-5.2
- Qwen3.6 35B-A3B
- Qwen3.6 27B
- MiniMax-M2.7
- Ling 2.6 1T
- Hy3 preview
- GLM-5.1
- DeepSeek-V4-Pro
- DeepSeek-V4-Flash
- Nemotron 3 Super
- Step 3.5 Flash
- Qwen3.5 397B-A17B
- MiniMax-M2.5
- GLM-5
- GLM-4.7
- DeepSeek-V3.2
- INTELLECT-3
- MiniMax-M2
- Qwen3-Next 80B-A3B
- GLM-4.5-Air
- GLM-4.5
- DeepSeek-R1
- DeepSeek-V3
Maths
If the draft at depth is accepted with probability given that the ones before it were, a step with drafted tokens emits
At batch 1 a decode step is memory-bound: its time is about the bytes it reads, (active weights plus the KV cache) plus modules of weights. The speed-up over plain decoding is
with bytes per weight. One module is a block's mixer and active FFN, the projection that merges the state with the next token's embedding, and its norms.
Code
// src/lib/chapters/model.ts — the memory-bound speed-up (excerpt)
const main = decodeBytes(spec, context, weightBytes, kvBytes).total;
const extra = depth * mtpModuleActive(spec) * weightBytes;
const tokens = mtpExpectedTokens(acceptance, depth);
const step = main + extra;