Notation¶
One symbol table for the whole of Foundations. Every chapter uses these and only these. If a chapter needs a new symbol, it is added here first.
Model shape¶
| Symbol | Name | Meaning | Llama-3.1-8B |
|---|---|---|---|
| \(d\) | hidden size | Width of the residual stream | 4096 |
| \(L\) | layers | Number of decoder blocks | 32 |
| \(H\) | attention heads | Query heads per layer | 32 |
| \(H_{kv}\) | KV heads | Key/value heads per layer (= \(H\) for MHA, 1 for MQA) | 8 |
| \(d_h\) | head dimension | Width of one head; usually \(d/H\) | 128 |
| \(d_{ff}\) | MLP intermediate | Width of the MLP hidden layer | 14336 |
| \(V\) | vocabulary | Number of token IDs | 128256 |
| \(N\) | parameters | Total weight count | 8.03 B |
Workload shape¶
| Symbol | Name | Meaning |
|---|---|---|
| \(B\) | batch size | Sequences processed together |
| \(T\) | context length | Tokens currently in the KV cache for one sequence |
| \(T_{in}\) | prompt length | Tokens in the input, processed during prefill |
| \(T_{out}\) | generation length | Tokens produced during decode |
| \(C\) | concurrency | Sequences in flight at the same time |
Machine¶
| Symbol | Name | Unit | A100 80 GB (SXM) |
|---|---|---|---|
| \(P_{peak}\) | peak compute | FLOP/s | 312 × 10¹² (dense bf16) |
| \(B_{mem}\) | memory bandwidth | byte/s | 2.039 × 10¹² (spec) |
| \(I\) | arithmetic intensity | FLOP/byte | — |
| \(I_{ridge}\) | ridge point | FLOP/byte | ≈ 153 |
| \(s\) | bytes per element | byte | 2 (bf16), 1 (int8), 0.5 (int4) |
Cost and latency¶
| Symbol | Name | Meaning |
|---|---|---|
| \(M_w\) | weight memory | \(N \cdot s\) bytes |
| \(M_{kv}\) | KV cache memory | \(2 L B T H_{kv} d_h s\) bytes |
| TTFT | time to first token | Queue + prefill + first sample |
| TPOT | time per output token | Steady-state decode step time |
| ITL | inter-token latency | Observed gap between streamed tokens |
| \(\lambda\) | arrival rate | Requests per second |
Conventions¶
- A FLOP is one operation. A fused multiply-add counts as two. A matrix multiply of \((m \times k)\) by \((k \times n)\) costs \(2mkn\) FLOPs.
- Bytes moved always names a boundary. Unless stated otherwise, the boundary is HBM. Intensity computed against L2 or SRAM is a different number and must be labelled.
- GiB vs GB. Memory capacity uses GiB (\(2^{30}\)). Bandwidth and vendor sheets use GB/s (\(10^9\)). Chapters state which. Mixing them is a 7.4% error and it compounds.
- Peak compute is dense, non-sparse, and dtype-matched. Comparing bf16 work against a sparse or FP8 peak is the most common way to fabricate a bad utilisation number.
The running example¶
Two models carry the entire book so the numbers compound instead of resetting.
The spine — Llama-3.1-8B. Dense, GQA, 8 B parameters. Every formula is instantiated on it.
{ "hidden_size": 4096, "num_hidden_layers": 32, "num_attention_heads": 32,
"num_key_value_heads": 8, "head_dim": 128, "intermediate_size": 14336,
"vocab_size": 128256, "max_position_embeddings": 131072, "rope_theta": 500000.0 }
Three numbers to memorise, because they recur in every chapter:
| Quantity | Value | Where it comes from |
|---|---|---|
| Weights in bf16 | 14.96 GiB | \(8.03\text{e}9 \times 2\) bytes |
| KV cache per token | 128 KiB | \(2 \times 32 \times 8 \times 128 \times 2\) bytes |
| KV cache per 8 K sequence | 1 GiB | \(8192 \times 128\) KiB |
The contrast — Mixtral-8x7B. Sparse MoE, 46.7 B total, 12.9 B active. Used whenever the point is that total and active parameters answer different questions.
The runnable — a 0.1–0.6 B toy. Small enough that every code artifact finishes in seconds on one GPU, or on CPU with the documented scale-down flag.