Encoder-decoder and the causal encoder-decoder
T5's split, and DeepSeek-V4.1-Flash's encoder-only prefill, simulated live.
Concept
Encoder-decoder was the original Transformer's shape, and T5's (Raffel et al., 2019): an encoder reads the whole input with bidirectional attention, and a decoder generates the output, attending causally to what it has written and, through cross-attention in every layer, to the encoder's output. Decoder-only models dropped the split: one causal stack reads the prompt and writes the answer. In this data set the original Transformer, T5 11B and Switch-C are encoder-decoders and BERT an encoder; every other model is a decoder, except one.
DeepSeek-V4.1-Flash brings a split back, in a different form, for a different reason. Its causal encoder-decoder (CED) divides its 40 layers into a 20-layer causal encoder and a 20-layer decoder, and projects the decoder's global keys and values from the last encoder layer's hidden state (section 2.2, eq. 1, p. 9). So a prompt token needs only the encoder, plus the decoder's key/value projections: prefill does about half the work (the paper: it "nearly halves prefill computation", p. 8; 8B parameters activated per token at prefill against 16B at decode). Decode still runs the whole model. The decoder's sliding-window layers need their own recent keys and values, so the last prompt tokens are replayed through the decoder, an approximation the paper simulates during post-training (section 3.2.2, p. 20). The idea, the paper says, was inspired by YOCO (Sun et al., 2024), whose "self-decoder" builds a cache that a "cross-decoder" reuses.
First, what the split saves on the real model, by this site's cost model. At an 8K-token prompt, CED's prefill is 0.458 of the same stack run as a decoder, and at 1M tokens 0.337: below one half, because in this model's recorded layout the decoder half's compressed attention keeps every entry (ratio 1) where the encoder half keeps one in two (ratio 2), so the half that prefill skips is the more expensive one.
Loading the interactive…
Concept
Then, what it does to a serving cluster. Prefill and decode already run on separate GPU pools in disaggregated serving. If a prompt token costs half as much, the cluster needs fewer prefill GPUs for the same load, and the freed GPUs can decode. Brief 11 added a CED option to Disaggregated_Inference_Sim to measure that, with two placements of the 128-token replay: on the prefill instance (the paper's), or as the decode instance's first step (SGLang RFC #39963's, where a prefill instance holds only the encoder).
The interactive below runs that simulator's own JavaScript engine, vendored byte for byte at the commit that added CED, in your browser. Its proxy is a dense Llama-3-70B shape split 40 + 40. One prompt token of the CED proxy touches 34.90B parameters against 69.50B for the decoder-only model (0.502 of them); at 8,192 tokens one prefill step takes 288.3 ms with the replay on the prefill instance and 283.5 ms encoder-only, against 564.3 ms. An encoder-only prefill instance also holds half the weights: on two H100s it has room for 220,048 tokens of cache, against 8,835 for the whole model.
The capacity search finds, for each split of six 4×H100 instances into
prefill and decode pools, the highest load at which 90% of requests meet
both SLOs. Press Re-measure live and your browser reruns every
bisection, checking each rate against results.md. With prompts four times
the outputs (2,048 : 512) the best decoder-only split is 4P2D at 25.72
req/s, and CED's moves to 3P3D, at 39.20 req/s with the replay on prefill
and 41.07 on decode: 1.52× and 1.60×. At 4,096 : 256 the gains are 2.06×
and 2.13×. Gains above 2× are queueing, not FLOPs: with a fixed TTFT SLO,
halving the prefill service time cuts the queueing delay by more than half.
Loading the simulator…
Encoder-decoder: 3 models in the data set
A bidirectional encoder and a decoder with cross-attention. Newest first; each links to its sourced page.
Causal encoder-decoder: 1 model in the data set
Prefill runs only the first half of the stack (DeepSeek V4.1). Newest first; each links to its sourced page.
Cross-layer KV sharing: 3 models in the data set
Some layers reuse another layer's keys and values. Newest first; each links to its sourced page.
Maths
With layers split into encoder and decoder layers, a prompt of tokens costs, in matrix-multiplication FLOPs,
against for the decoder-only stack, where is a decoder layer's key/value projection and the replay window. For and the ratio tends to (ignoring the output head). For the 70B proxy, section 16 counts 34.90B parameters per prompt token (the encoder plus 0.67B of decoder K/V projections) against 69.50B per decode token, output head included: 0.502. The paper writes the same saving as prefill going from to (section 2.2, p. 9).
Code
// src/lib/disagg/engine.ts — the capacity search, as the simulator's search.py
while (att(hi) >= target) {
runs += 1;
lo = hi;
hi = hi * 1.5;
}
runs += 1;
while (hi - lo > relTol * hi) {
const mid = (lo + hi) / 2;
runs += 1;
if (att(mid) >= target) lo = mid;
else hi = mid;
}
tests/unit/disagg/parity.test.ts runs the simulator's own ten CED parity
configurations against Python fixtures from the same commit (every request's
timestamps, every instance's energy); tests/unit/disagg/search.test.ts
reruns all 72 bisections of sections 17 and 18 and requires every rate to
equal the recorded one exactly. The workloads are Python's own draws,
stored as unit exponentials so that any rate rebuilds Python's arrivals bit
for bit.