Skip to content

Notation

One symbol table for the whole of Foundations. Every chapter uses these and only these. If a chapter needs a new symbol, it is added here first.


Model shape

Symbol Name Meaning Llama-3.1-8B
\(d\) hidden size Width of the residual stream 4096
\(L\) layers Number of decoder blocks 32
\(H\) attention heads Query heads per layer 32
\(H_{kv}\) KV heads Key/value heads per layer (= \(H\) for MHA, 1 for MQA) 8
\(d_h\) head dimension Width of one head; usually \(d/H\) 128
\(d_{ff}\) MLP intermediate Width of the MLP hidden layer 14336
\(V\) vocabulary Number of token IDs 128256
\(N\) parameters Total weight count 8.03 B

Workload shape

Symbol Name Meaning
\(B\) batch size Sequences processed together
\(T\) context length Tokens currently in the KV cache for one sequence
\(T_{in}\) prompt length Tokens in the input, processed during prefill
\(T_{out}\) generation length Tokens produced during decode
\(C\) concurrency Sequences in flight at the same time

Machine

Symbol Name Unit A100 80 GB (SXM)
\(P_{peak}\) peak compute FLOP/s 312 × 10¹² (dense bf16)
\(B_{mem}\) memory bandwidth byte/s 2.039 × 10¹² (spec)
\(I\) arithmetic intensity FLOP/byte —
\(I_{ridge}\) ridge point FLOP/byte ≈ 153
\(s\) bytes per element byte 2 (bf16), 1 (int8), 0.5 (int4)

Cost and latency

Symbol Name Meaning
\(M_w\) weight memory \(N \cdot s\) bytes
\(M_{kv}\) KV cache memory \(2 L B T H_{kv} d_h s\) bytes
TTFT time to first token Queue + prefill + first sample
TPOT time per output token Steady-state decode step time
ITL inter-token latency Observed gap between streamed tokens
\(\lambda\) arrival rate Requests per second

Conventions

  • A FLOP is one operation. A fused multiply-add counts as two. A matrix multiply of \((m \times k)\) by \((k \times n)\) costs \(2mkn\) FLOPs.
  • Bytes moved always names a boundary. Unless stated otherwise, the boundary is HBM. Intensity computed against L2 or SRAM is a different number and must be labelled.
  • GiB vs GB. Memory capacity uses GiB (\(2^{30}\)). Bandwidth and vendor sheets use GB/s (\(10^9\)). Chapters state which. Mixing them is a 7.4% error and it compounds.
  • Peak compute is dense, non-sparse, and dtype-matched. Comparing bf16 work against a sparse or FP8 peak is the most common way to fabricate a bad utilisation number.

The running example

Two models carry the entire book so the numbers compound instead of resetting.

The spine — Llama-3.1-8B. Dense, GQA, 8 B parameters. Every formula is instantiated on it.

{ "hidden_size": 4096, "num_hidden_layers": 32, "num_attention_heads": 32,
  "num_key_value_heads": 8, "head_dim": 128, "intermediate_size": 14336,
  "vocab_size": 128256, "max_position_embeddings": 131072, "rope_theta": 500000.0 }

Three numbers to memorise, because they recur in every chapter:

Quantity Value Where it comes from
Weights in bf16 14.96 GiB \(8.03\text{e}9 \times 2\) bytes
KV cache per token 128 KiB \(2 \times 32 \times 8 \times 128 \times 2\) bytes
KV cache per 8 K sequence 1 GiB \(8192 \times 128\) KiB

The contrast — Mixtral-8x7B. Sparse MoE, 46.7 B total, 12.9 B active. Used whenever the point is that total and active parameters answer different questions.

The runnable — a 0.1–0.6 B toy. Small enough that every code artifact finishes in seconds on one GPU, or on CPU with the documented scale-down flag.