T2.4 — Prefill and Decode Are Different Workloads¶
In one line: Prefill processes many prompt tokens with reusable matrix work, while decode processes one new token per sequence and repeatedly streams weights and cache state.
| Skill | 00-foundations |
| Module | M2 — Transformer Internals from the Inference Angle |
| Audience | New graduate engineer |
| Time | ~110 min |
| Prereqs | T2.1–T2.3; M1T2 |
| Status | built |
Why this matters¶
Time to first token and time per output token come from different phases with different bottlenecks. Treating them as one forward pass leads to bad batching, misleading benchmarks, and hardware choices that optimize the wrong part of the request.
The mental model¶
An inference request has two phases. Prefill reads the submitted prompt and constructs the initial state. Decode repeatedly uses that state to produce one new token per active sequence. They use the same model weights, but they present different matrix shapes to the GPU.

Prefill consumes the prompt in parallel; decode advances one generated step at a time.
Think of prefill as grading a stack of pages with one answer key: the key is loaded and reused across many rows. Decode at low batch is grading one new line at a time, so the same large answer key is repeatedly fetched for little work.
sequenceDiagram
participant U as Request
participant P as Prefill
participant C as KV cache
participant D as Decode loop
U->>P: T prompt tokens
P->>C: write T K/V entries
P-->>U: first-token logits
loop one step per output token
D->>C: read history, append one entry
D-->>U: one token
end
Prefill creates the cache and dominates time to first token; decode extends it and dominates inter-token latency.
The mechanism¶
For a linear projection with weight shape d × d, processing m token rows requires about 2md² FLOPs. The weight matrix contains d² elements. If weight traffic dominates and is read once, increasing m performs more work with the same weight bytes.
During prefill, m can be hundreds or thousands of prompt tokens across the batch. The projection resembles a matrix–matrix multiplication and has substantial reuse. During low-batch decode, m is the number of active sequences—possibly one—and the operation resembles a matrix–vector multiplication. It has far less arithmetic per weight byte.

Prefill reuses one weight load across many rows; low-batch decode does little work per load.
The artifact isolates this effect with one projection. With d=2048 on the local CPU diagnostic, 256 prefill rows have an optimistic intensity of 102.4 FLOP/byte, while eight decode rows have 3.97 FLOP/byte. These are analytical minimum-traffic values around a measured operation; they do not claim GPU HBM traffic. Their ratio shows why identical weights can create different hardware regimes.
Attention work differs too¶
Prefill attention computes interactions among prompt positions under a causal mask. A naive implementation materializes a T × T score matrix per head, although memory-efficient attention kernels tile the calculation without storing that full matrix in HBM. Arithmetic still grows roughly with prompt length squared for full attention.
Decode has one new query per active sequence and compares it with all cached keys. Per step, attention work and KV traffic grow roughly linearly with current context length. Across a long generated continuation, those steps accumulate.
The latency metrics follow the phases¶
Time to first token (TTFT) includes queueing, prompt preparation, prefill, and the first sampling decision. Longer prompts usually increase prefill work and TTFT. Time per output token (TPOT) or inter-token latency (ITL) measures decode cadence after generation begins. It is influenced by batch size, current context, cache traffic, scheduling, and sampling.
End-to-end latency combines both plus application overhead. Reporting only total latency hides whether a change helped prompt processing or token streaming. A configuration can reduce TTFT while worsening TPOT, or improve total throughput by batching while making each user wait longer.
Consider a 2,000-token prompt followed by 100 output tokens. If prefill takes 300 ms and decode takes 20 ms per token, model time is roughly 300 + 100×20 = 2,300 ms. Halving prefill saves 150 ms; halving TPOT saves 1,000 ms. For a request expecting one output token, the priority reverses. Workload shape determines which phase matters most.
Why batching affects them differently¶
Prefill already contains many token rows, so additional requests may improve shape utilization but do not create the same dramatic weight-reuse change. Decode at batch one uses each weight for one row. At batch 16, the same weight can serve 16 active sequences, raising optimistic weight arithmetic intensity toward 16 FLOP/byte.
This is not free latency. Requests may wait in a queue so the scheduler can form work, and a large batch can become compute-bound. Continuous batching tries to keep the GPU productive by inserting new work as sequences finish rather than forcing an entire static batch to wait for the longest response.
Batch is not only request count. Engines often schedule a budget of batched tokens: a prefill chunk may contribute hundreds of rows while each decode sequence contributes one. The scheduler chooses a mixture that fits this budget. Record sequence limits, token budget, and chunking settings because they determine the shapes reaching the kernels.
Why disaggregation appears¶
Prefill favors compute-dense prompt work; decode favors steady low-latency execution with a growing cache. Mixing them on one worker can cause interference: a long prefill delays decode steps, while strict decode scheduling can fragment prefill. Disaggregated architectures place phases on different workers and transfer KV state between them. That separation helps only if scheduling and cache-transfer cost are lower than the interference removed.

Separating phases removes interference only if KV transfer costs less than the contention avoided.
In practice¶
Benchmark prompt and output length distributions separately. Report TTFT percentiles, TPOT/ITL percentiles, total output tokens per second, and goodput under an explicit SLO. An average “request latency” is insufficient.
Shape the load realistically. A single fixed prompt length hides chunked prefill, cache pressure, and scheduler interaction. Record batch-token limits and active sequence counts because they determine the matrix rows actually reaching kernels.
Use profiler traces to label prefill and decode regions. High tensor-core activity during prefill and high HBM traffic during low-batch decode support the cost model; low utilization everywhere instead suggests small shapes, launch gaps, throttling, or software overhead.
Avoid comparing phases with different features silently enabled. Prefix caching can skip prefill, speculative decoding changes decode work, and quantization can affect phases differently. Warm up compilation and CUDA graphs before steady-state measurement, while measuring cold behavior separately if real requests experience it.
Failure modes¶
| Symptom | Cause | Fix |
|---|---|---|
| Throughput rises while user latency worsens | Larger batches improve reuse but add queueing or per-step work | Plot latency–throughput tradeoffs and enforce an SLO |
| TTFT spikes when another request arrives | Large prefills block decode on a shared worker | Use chunked prefill, scheduling limits, or evaluate disaggregation |
| Decode slows as conversation grows | Every new query reads a longer KV history | Measure by context bucket and inspect cache bandwidth |
| Prefill is unexpectedly slow | Attention backend, shape, compilation, or tokenization dominates | Profile the phase rather than assuming compute saturation |
| CPU diagnostic is presented as GPU evidence | Timing environment does not represent tensor cores or HBM | Rerun on CUDA and retain the diagnostic label until then |
Do it¶
Run the phase-shape artifact. Sweep prompt rows and decode batch, calculate optimistic intensity before each run, and identify where measured execution changes character on your GPU.
Success means explaining TTFT and TPOT with separate evidence. Do not claim an end-to-end serving speedup from this isolated projection alone.
Check¶
- Why can the same projection be compute-friendly during prefill and bandwidth-limited during decode?
- Which phase primarily contributes to TTFT, and which determines token streaming cadence?
- Why can increasing decode batch improve throughput while worsening per-user latency?