Skill 00 — Foundations¶
The physics under everything else. Not a course in transformers — a course in why LLM inference costs what it costs.
| Duration | 2 weeks (~25 hrs) |
| Impact | ★★★☆☆ — low market value alone, but every later skill collapses without it |
| Prerequisites | Python, basic linear algebra, comfort with PyTorch tensors |
| Hardware | 1 GPU. Toy models only (0.1–0.6 B) |
| Status | planned |
Outcome¶
By the end you can, on a whiteboard, without looking anything up:
- Compute VRAM for weights, KV cache, and activations for any model config
- Explain why prefill is compute-bound and decode is memory-bandwidth-bound — and derive it
- Predict whether a given change will help throughput, latency, both, or neither
- Read a
config.jsonfrom Hugging Face and immediately state the serving implications - Explain GQA/MQA/MLA as KV-cache decisions, not as architecture trivia
The skip test¶
If you can answer all five right now, skip to Skill 01. Otherwise, do the two weeks.
- A 7B model, GQA with 8 KV heads, 32 layers, head_dim 128, bf16. How much KV cache for 100 concurrent 4k-token sequences?
- Why does batch size improve decode throughput almost for free, but barely help prefill?
- What is arithmetic intensity, and what is it for a decode step at batch size 1?
- Why is int4 weight-only quantization a latency win but not a throughput win at high batch?
- What breaks first when you double context length — compute, memory, or bandwidth?
Modules¶
| M | Module | Hrs | Theme |
|---|---|---|---|
| M1 | GPU & Systems Literacy | 8 | The hardware you're actually programming |
| M2 | Transformer Internals from the Inference Angle | 10 | Where the FLOPs and the bytes go |
| M3 | The Cost Model | 7 | Turn architecture into dollars |
M1 — GPU & Systems Literacy¶
Status: draft — all five articles and runnable artifacts are authored; CUDA/NCCL evidence is pending on the target rig.
| ID | Topic | Key concepts | Code artifact | Visual | Evidence |
|---|---|---|---|---|---|
| M1T1 | The memory hierarchy that matters | HBM vs L2 vs SRAM, capacity vs bandwidth, why registers are free and HBM is not | Micro-benchmark: measured HBM bandwidth vs spec sheet | Gen: tiered pyramid (Pattern D) | Your GPU's real achievable bandwidth, as a % of spec |
| M1T2 | Roofline, arithmetic intensity, MFU | FLOPs vs bytes moved, the ridge point, MFU vs HFU, why 40% MFU is good | Roofline plotter for a matmul sweep | Chart: roofline with matmul points overlaid | A roofline plot of your GPU with GEMM shapes marked |
| M1T3 | Precision formats and what they cost | fp32/tf32/bf16/fp16/fp8/int8/int4/nvfp4, exponent vs mantissa, dynamic range, why bf16 won | Numeric-range explorer: quantize a weight tensor to each format, plot error | Gen: bit-layout comparison (Pattern B) | Error-vs-format table for a real weight tensor |
| M1T4 | Multi-GPU plumbing | NVLink vs PCIe, NCCL collectives (all-reduce/all-gather/reduce-scatter), latency vs bandwidth regimes | NCCL bandwidth test across 2 and 4 GPUs, NVLink vs forced-PCIe | Mermaid: collective operation diagrams | All-reduce bandwidth curve vs message size |
| M1T5 | The toolchain and its failure modes | driver/CUDA/toolkit/PyTorch compatibility, nvidia-smi, DCGM, ECC, clock throttling, CUDA_VISIBLE_DEVICES |
Diagnostic script: dump full environment + health | Mermaid: version-compatibility decision tree | A one-page health report for your own rig |
Failure modes to teach explicitly: silent clock throttling, ECC-disabled memory differences, fragmented VRAM from a zombie process, PyTorch built for the wrong CUDA minor.
M2 — Transformer Internals from the Inference Angle¶
Status: built — six new-graduate articles, runnable mechanism artifacts, evidence charts, exact diagrams, generated visuals, and LAB-M2 are present; real-checkpoint/GPU review remains pending.
| ID | Topic | Key concepts | Code artifact | Visual | Evidence |
|---|---|---|---|---|---|
| M2T1 | Anatomy of a decoder block | Embedding, attention, MLP, RMSNorm, residual stream; parameter count derivation ≈12·d²·L |
Parameter counter that reads any HF config.json and breaks down params by component |
Mermaid: block diagram with tensor shapes | Param breakdown for 3 real models, verified against the checkpoint |
| M2T2 | Attention and the KV cache | Why caching K,V is mandatory; cache size formula; growth with batch × seqlen; the memory wall | Attention + KV cache in ~100 lines of raw PyTorch, no libraries | Gen: cache growth as accumulating blocks (Pattern A) | Working single-file generation loop; measured cache size vs formula |
| M2T3 | MHA → GQA → MQA → MLA | KV head sharing as a compression decision; quality vs cache-size tradeoff; MLA's latent compression | Convert an MHA config to GQA variants, plot cache size vs KV heads | Gen: head-sharing contrast (Pattern B) | Cache-size-vs-quality tradeoff table |
| M2T4 | Prefill vs decode | Two workloads, opposite hardware profiles; why batching helps one and not the other; the seed of disaggregation | Instrumented forward pass: time and FLOPs for prefill vs each decode step | Gen: two-phase pipeline (Pattern C) | Measured FLOPs/byte for prefill vs decode on the same model |
| M2T5 | Positional encoding and long context | RoPE, theta scaling, YaRN/NTK extension, why context length is a memory problem before it's a quality problem | RoPE implemented from scratch; extend a model's context and observe degradation | Mermaid: rotation intuition + scaling methods | Perplexity vs context length curve past the trained window |
| M2T6 | MoE: what changes | Router, top-k, shared experts, sparse vs total params, why MoE is a memory-capacity trade for compute savings | Toy MoE layer; measure active vs total params and expert-load imbalance | Gen: routing illustration (Pattern C) | Expert-load histogram showing imbalance |
M3 — The Cost Model¶
| ID | Topic | Key concepts | Code artifact | Visual | Evidence |
|---|---|---|---|---|---|
| M3T1 | The VRAM budget | weights + KV + activations + fragmentation + CUDA context; why "80GB" is never 80GB | VRAM calculator: model config + batch + seqlen → full breakdown | Gen: budget as a filling container (Pattern A) | Calculator predictions within 10% of measured nvidia-smi |
| M3T2 | The latency budget | TTFT decomposition (queue + prefill + first token), TPOT/ITL, why p99 ≠ p50 × 2 | Latency decomposer: instrument each stage of a request | Mermaid: sequence diagram of request lifecycle | A stacked-bar latency breakdown for one request |
| M3T3 | The throughput ceiling | Bandwidth-bound decode ceiling tok/s ≈ BW / bytes_per_param_read; where batching saturates |
Theoretical-vs-measured ceiling calculator | Chart: theoretical ceiling vs measured, vs batch size | The gap between theory and reality, explained |
| M3T4 | $/1M tokens | GPU-hour cost → cost per million input and output tokens at a given SLO; utilization as the hidden multiplier | Cost model spreadsheet/notebook, parameterized | Chart: $/1M tokens vs concurrency, with SLO cutoff | Your own cost curve — the artifact you show leadership |
Labs¶
| Lab | Goal | Success criterion |
|---|---|---|
| LAB-M1 | Profile your own rig | Health report + measured bandwidth within 15% of a known-good reference |
| LAB-M2 | Build a generation loop from scratch | Generates coherent text with a working KV cache; matches HF generate() token-for-token at temperature 0 |
| LAB-M3 | Predict-then-measure | Predict VRAM and throughput for 3 unseen configs; all three within 15% before you're allowed to proceed |
LAB-M3 is the gate. It's the difference between having read this and knowing it.
Capstone¶
Deliverable: llm-sizing-tool — a small CLI + notebook that takes a Hugging Face model ID, a workload description (input/output length distribution, concurrency, SLO), and hardware spec, and outputs:
- VRAM breakdown and whether it fits
- Recommended TP degree and max batch size
- Predicted TTFT, TPOT, and throughput
- $/1M input and output tokens
- The binding constraint, named ("you are KV-cache-capacity bound")
Rubric: predictions within 20% of measured reality on three models you haven't tested yet. Everything in Skill 01 will be used to falsify and refine this tool — that's the point.
Assessment¶
| Tier | Count | Example |
|---|---|---|
| Recall | 7 | State the KV cache size formula and define every term |
| Apply | 12 | Given this config.json and 60 concurrent 8k-token users, does it fit on one A100 80GB? Show working |
| Design | 4 | Your model is KV-bound. Rank five interventions by expected impact and justify the ordering |
Practical challenge: given an unfamiliar model config and a hardware spec, produce a complete sizing recommendation in 30 minutes, then verify it on real hardware.
Asset inventory¶
| Type | Count |
|---|---|
| Theory files | 15 |
| Code artifacts | 15 |
| Mermaid diagrams | ~20 |
| Generated images | ~10 |
| Charts | ~12 |
| Labs | 3 |
| Capstone | 1 |
Primary sources¶
- Karpathy — Let's build GPT +
nanoGPT - vLLM PagedAttention paper (arXiv 2309.06180) — read §2 for the memory analysis
- Efficient Memory Management for LLM Serving + the FlashAttention papers (for the SRAM/HBM argument)
- GPU MODE lecture series (roofline, memory hierarchy)
- Hugging Face
transformerssource:modeling_llama.py— read the actual attention implementation - NVIDIA A100 architecture whitepaper (for the hardware numbers you'll benchmark against)