Glossary¶
Terms defined where they are first introduced, collected here for lookup. The chapter link is the place the term is taught, not merely mentioned.
A¶
Activation memory — Transient tensors produced during a forward pass. Small during decode, large during prefill because it scales with the number of tokens processed at once. Chapter 3.1.
All-reduce — A collective in which every rank contributes a tensor and every rank ends up with the elementwise sum. The workhorse of tensor parallelism. Chapter 1.4
Arithmetic intensity (\(I\)) — FLOPs performed per byte moved across a named memory boundary. Determines which side of the roofline ridge an operation lands on. Chapter 1.2
B¶
Bandwidth-bound — Limited by how fast bytes arrive from HBM, not by arithmetic. Decode is bandwidth-bound at low batch size. Also called memory-bound. Chapter 1.2
Block (KV) — A fixed-size unit of KV cache allocation, typically 16 tokens. Paging KV in blocks removes the fragmentation caused by reserving a contiguous maximum-length buffer. Chapter 2.2
C¶
Coalesced access — Threads within a warp reading contiguous addresses, so the memory system services them with the minimum number of transactions. Uncoalesced access can cost 8–32× more. Chapter 1.0
Compute-bound — Limited by arithmetic throughput. Prefill on a large prompt is compute-bound. Chapter 1.2
CUDA graph — A captured, replayable sequence of kernel launches. Removes per-launch CPU overhead, which matters when a decode step is a few hundred microseconds long. Chapter 1.0
D¶
Decode — The autoregressive phase: one token in, one token out, per step, per sequence. Memory-bandwidth-bound at low batch size. Chapter 2.4
Disaggregation — Running prefill and decode on separate GPU pools because they have opposite hardware profiles and opposite SLOs. Chapter 2.4
E¶
ECC — Error-correcting memory. Enabled by default on datacentre GPUs; costs a few percent of capacity and bandwidth. Disabling it changes your baseline. Chapter 1.5
Expert — One MLP inside a Mixture-of-Experts layer. Only the top-\(k\) selected experts run per token, but all of them occupy VRAM. Chapter 2.6
G¶
GQA (grouped-query attention) — Several query heads share one key/value head. Reduces KV cache size by \(H/H_{kv}\) with modest quality cost. Llama-3.1-8B uses 32 query heads over 8 KV heads. Chapter 2.3
Goodput — Requests per second that actually meet their SLO. Throughput counts everything; goodput counts only what you can sell. Chapter 3.4.
H¶
HBM — High-bandwidth memory: the GPU's main memory. Large (40–80 GB), fast in absolute terms (2 TB/s), and desperately slow relative to the arithmetic units it feeds. Chapter 1.1
HFU (hardware FLOPs utilisation) — Fraction of peak achieved counting every executed FLOP, including recomputation. Can exceed MFU without the job getting faster. Chapter 1.2
K¶
KV cache — Stored keys and values for every previous token, per layer. Turns decode from quadratic re-projection into a single new projection per step, at the price of memory that grows linearly with context and concurrency. Chapter 2.2
M¶
MFU (model FLOPs utilisation) — Useful model FLOPs per second divided by dtype-matched peak. 40% is good for an end-to-end workload; 40% for one large aligned GEMM is a bug. Chapter 1.2
MLA (multi-head latent attention) — Compresses K and V into a shared low-rank latent that is cached instead of full K/V tensors. A different compression axis from GQA. Chapter 2.3
N¶
NVLink — High-bandwidth GPU-to-GPU interconnect. On A100, 600 GB/s bidirectional versus 64 GB/s for PCIe Gen4 ×16. The difference decides whether tensor parallelism is viable. Chapter 1.4
O¶
Occupancy — Active warps per SM as a fraction of the hardware maximum. Enough occupancy hides memory latency; more than enough buys nothing. Chapter 1.0
P¶
Prefill — Processing the whole prompt in one forward pass to populate the KV cache and emit the first token. Compute-bound and the dominant term in TTFT. Chapter 2.4
R¶
Residual stream — The \(B \times T \times d\) tensor that every block reads from and adds back into. The transformer's shared workspace. Chapter 2.1
Ridge point (\(I_{ridge}\)) — \(P_{peak} / B_{mem}\). The arithmetic intensity below which peak compute is unreachable in principle. ≈153 FLOP/byte on an A100. Chapter 1.2
RoPE — Rotary position embedding. Encodes position by rotating query and key vectors, so attention scores depend on relative displacement. Chapter 2.5
S¶
SM (streaming multiprocessor) — The GPU's independent processor core. An A100 has 108. Chapter 1.0
SRAM (shared memory / L1) — Fast on-chip scratchpad, ~192 KB per SM. FlashAttention exists to keep attention working inside it instead of round-tripping to HBM. Chapter 1.1
T¶
Tensor core — Fixed-function matrix-multiply unit. Delivers the headline TFLOP/s number, and only when shapes, alignment, and dtype cooperate. Chapter 1.0
TPOT (time per output token) — Steady-state decode step latency. Its reciprocal is the per-user streaming rate. Chapter 3.2.
TTFT (time to first token) — Queue wait + prefill + first sample. What users perceive as responsiveness. Chapter 3.2.
W¶
Warp — 32 threads executing in lockstep. The real unit of GPU scheduling. Chapter 1.0
W4A16 / W8A8 — Quantisation shorthand: weight bits / activation bits. W4A16 is a latency win at low batch and often a throughput loss at high batch. Chapter 1.3