Skip to content

Glossary

Terms defined where they are first introduced, collected here for lookup. The chapter link is the place the term is taught, not merely mentioned.


A

Activation memory — Transient tensors produced during a forward pass. Small during decode, large during prefill because it scales with the number of tokens processed at once. Chapter 3.1.

All-reduce — A collective in which every rank contributes a tensor and every rank ends up with the elementwise sum. The workhorse of tensor parallelism. Chapter 1.4

Arithmetic intensity (\(I\)) — FLOPs performed per byte moved across a named memory boundary. Determines which side of the roofline ridge an operation lands on. Chapter 1.2

B

Bandwidth-bound — Limited by how fast bytes arrive from HBM, not by arithmetic. Decode is bandwidth-bound at low batch size. Also called memory-bound. Chapter 1.2

Block (KV) — A fixed-size unit of KV cache allocation, typically 16 tokens. Paging KV in blocks removes the fragmentation caused by reserving a contiguous maximum-length buffer. Chapter 2.2

C

Coalesced access — Threads within a warp reading contiguous addresses, so the memory system services them with the minimum number of transactions. Uncoalesced access can cost 8–32× more. Chapter 1.0

Compute-bound — Limited by arithmetic throughput. Prefill on a large prompt is compute-bound. Chapter 1.2

CUDA graph — A captured, replayable sequence of kernel launches. Removes per-launch CPU overhead, which matters when a decode step is a few hundred microseconds long. Chapter 1.0

D

Decode — The autoregressive phase: one token in, one token out, per step, per sequence. Memory-bandwidth-bound at low batch size. Chapter 2.4

Disaggregation — Running prefill and decode on separate GPU pools because they have opposite hardware profiles and opposite SLOs. Chapter 2.4

E

ECC — Error-correcting memory. Enabled by default on datacentre GPUs; costs a few percent of capacity and bandwidth. Disabling it changes your baseline. Chapter 1.5

Expert — One MLP inside a Mixture-of-Experts layer. Only the top-\(k\) selected experts run per token, but all of them occupy VRAM. Chapter 2.6

G

GQA (grouped-query attention) — Several query heads share one key/value head. Reduces KV cache size by \(H/H_{kv}\) with modest quality cost. Llama-3.1-8B uses 32 query heads over 8 KV heads. Chapter 2.3

Goodput — Requests per second that actually meet their SLO. Throughput counts everything; goodput counts only what you can sell. Chapter 3.4.

H

HBM — High-bandwidth memory: the GPU's main memory. Large (40–80 GB), fast in absolute terms (2 TB/s), and desperately slow relative to the arithmetic units it feeds. Chapter 1.1

HFU (hardware FLOPs utilisation) — Fraction of peak achieved counting every executed FLOP, including recomputation. Can exceed MFU without the job getting faster. Chapter 1.2

K

KV cache — Stored keys and values for every previous token, per layer. Turns decode from quadratic re-projection into a single new projection per step, at the price of memory that grows linearly with context and concurrency. Chapter 2.2

M

MFU (model FLOPs utilisation) — Useful model FLOPs per second divided by dtype-matched peak. 40% is good for an end-to-end workload; 40% for one large aligned GEMM is a bug. Chapter 1.2

MLA (multi-head latent attention) — Compresses K and V into a shared low-rank latent that is cached instead of full K/V tensors. A different compression axis from GQA. Chapter 2.3

N

NVLink — High-bandwidth GPU-to-GPU interconnect. On A100, 600 GB/s bidirectional versus 64 GB/s for PCIe Gen4 ×16. The difference decides whether tensor parallelism is viable. Chapter 1.4

O

Occupancy — Active warps per SM as a fraction of the hardware maximum. Enough occupancy hides memory latency; more than enough buys nothing. Chapter 1.0

P

Prefill — Processing the whole prompt in one forward pass to populate the KV cache and emit the first token. Compute-bound and the dominant term in TTFT. Chapter 2.4

R

Residual stream — The \(B \times T \times d\) tensor that every block reads from and adds back into. The transformer's shared workspace. Chapter 2.1

Ridge point (\(I_{ridge}\)) — \(P_{peak} / B_{mem}\). The arithmetic intensity below which peak compute is unreachable in principle. ≈153 FLOP/byte on an A100. Chapter 1.2

RoPE — Rotary position embedding. Encodes position by rotating query and key vectors, so attention scores depend on relative displacement. Chapter 2.5

S

SM (streaming multiprocessor) — The GPU's independent processor core. An A100 has 108. Chapter 1.0

SRAM (shared memory / L1) — Fast on-chip scratchpad, ~192 KB per SM. FlashAttention exists to keep attention working inside it instead of round-tripping to HBM. Chapter 1.1

T

Tensor core — Fixed-function matrix-multiply unit. Delivers the headline TFLOP/s number, and only when shapes, alignment, and dtype cooperate. Chapter 1.0

TPOT (time per output token) — Steady-state decode step latency. Its reciprocal is the per-user streaming rate. Chapter 3.2.

TTFT (time to first token) — Queue wait + prefill + first sample. What users perceive as responsiveness. Chapter 3.2.

W

Warp — 32 threads executing in lockstep. The real unit of GPU scheduling. Chapter 1.0

W4A16 / W8A8 — Quantisation shorthand: weight bits / activation bits. W4A16 is a latency win at low batch and often a throughput loss at high batch. Chapter 1.3