Skip to content

Skill 00 — Foundations

The physics under everything else. Not a course in transformers — a course in why LLM inference costs what it costs.

Duration 2 weeks (~25 hrs)
Impact ★★★☆☆ — low market value alone, but every later skill collapses without it
Prerequisites Python, basic linear algebra, comfort with PyTorch tensors
Hardware 1 GPU. Toy models only (0.1–0.6 B)
Status planned

Outcome

By the end you can, on a whiteboard, without looking anything up:

  • Compute VRAM for weights, KV cache, and activations for any model config
  • Explain why prefill is compute-bound and decode is memory-bandwidth-bound — and derive it
  • Predict whether a given change will help throughput, latency, both, or neither
  • Read a config.json from Hugging Face and immediately state the serving implications
  • Explain GQA/MQA/MLA as KV-cache decisions, not as architecture trivia

The skip test

If you can answer all five right now, skip to Skill 01. Otherwise, do the two weeks.

  1. A 7B model, GQA with 8 KV heads, 32 layers, head_dim 128, bf16. How much KV cache for 100 concurrent 4k-token sequences?
  2. Why does batch size improve decode throughput almost for free, but barely help prefill?
  3. What is arithmetic intensity, and what is it for a decode step at batch size 1?
  4. Why is int4 weight-only quantization a latency win but not a throughput win at high batch?
  5. What breaks first when you double context length — compute, memory, or bandwidth?

Modules

M Module Hrs Theme
M1 GPU & Systems Literacy 8 The hardware you're actually programming
M2 Transformer Internals from the Inference Angle 10 Where the FLOPs and the bytes go
M3 The Cost Model 7 Turn architecture into dollars

M1 — GPU & Systems Literacy

Status: draft — all five articles and runnable artifacts are authored; CUDA/NCCL evidence is pending on the target rig.

ID Topic Key concepts Code artifact Visual Evidence
M1T1 The memory hierarchy that matters HBM vs L2 vs SRAM, capacity vs bandwidth, why registers are free and HBM is not Micro-benchmark: measured HBM bandwidth vs spec sheet Gen: tiered pyramid (Pattern D) Your GPU's real achievable bandwidth, as a % of spec
M1T2 Roofline, arithmetic intensity, MFU FLOPs vs bytes moved, the ridge point, MFU vs HFU, why 40% MFU is good Roofline plotter for a matmul sweep Chart: roofline with matmul points overlaid A roofline plot of your GPU with GEMM shapes marked
M1T3 Precision formats and what they cost fp32/tf32/bf16/fp16/fp8/int8/int4/nvfp4, exponent vs mantissa, dynamic range, why bf16 won Numeric-range explorer: quantize a weight tensor to each format, plot error Gen: bit-layout comparison (Pattern B) Error-vs-format table for a real weight tensor
M1T4 Multi-GPU plumbing NVLink vs PCIe, NCCL collectives (all-reduce/all-gather/reduce-scatter), latency vs bandwidth regimes NCCL bandwidth test across 2 and 4 GPUs, NVLink vs forced-PCIe Mermaid: collective operation diagrams All-reduce bandwidth curve vs message size
M1T5 The toolchain and its failure modes driver/CUDA/toolkit/PyTorch compatibility, nvidia-smi, DCGM, ECC, clock throttling, CUDA_VISIBLE_DEVICES Diagnostic script: dump full environment + health Mermaid: version-compatibility decision tree A one-page health report for your own rig

Failure modes to teach explicitly: silent clock throttling, ECC-disabled memory differences, fragmented VRAM from a zombie process, PyTorch built for the wrong CUDA minor.


M2 — Transformer Internals from the Inference Angle

Status: built — six new-graduate articles, runnable mechanism artifacts, evidence charts, exact diagrams, generated visuals, and LAB-M2 are present; real-checkpoint/GPU review remains pending.

ID Topic Key concepts Code artifact Visual Evidence
M2T1 Anatomy of a decoder block Embedding, attention, MLP, RMSNorm, residual stream; parameter count derivation ≈12·d²·L Parameter counter that reads any HF config.json and breaks down params by component Mermaid: block diagram with tensor shapes Param breakdown for 3 real models, verified against the checkpoint
M2T2 Attention and the KV cache Why caching K,V is mandatory; cache size formula; growth with batch × seqlen; the memory wall Attention + KV cache in ~100 lines of raw PyTorch, no libraries Gen: cache growth as accumulating blocks (Pattern A) Working single-file generation loop; measured cache size vs formula
M2T3 MHA → GQA → MQA → MLA KV head sharing as a compression decision; quality vs cache-size tradeoff; MLA's latent compression Convert an MHA config to GQA variants, plot cache size vs KV heads Gen: head-sharing contrast (Pattern B) Cache-size-vs-quality tradeoff table
M2T4 Prefill vs decode Two workloads, opposite hardware profiles; why batching helps one and not the other; the seed of disaggregation Instrumented forward pass: time and FLOPs for prefill vs each decode step Gen: two-phase pipeline (Pattern C) Measured FLOPs/byte for prefill vs decode on the same model
M2T5 Positional encoding and long context RoPE, theta scaling, YaRN/NTK extension, why context length is a memory problem before it's a quality problem RoPE implemented from scratch; extend a model's context and observe degradation Mermaid: rotation intuition + scaling methods Perplexity vs context length curve past the trained window
M2T6 MoE: what changes Router, top-k, shared experts, sparse vs total params, why MoE is a memory-capacity trade for compute savings Toy MoE layer; measure active vs total params and expert-load imbalance Gen: routing illustration (Pattern C) Expert-load histogram showing imbalance

M3 — The Cost Model

ID Topic Key concepts Code artifact Visual Evidence
M3T1 The VRAM budget weights + KV + activations + fragmentation + CUDA context; why "80GB" is never 80GB VRAM calculator: model config + batch + seqlen → full breakdown Gen: budget as a filling container (Pattern A) Calculator predictions within 10% of measured nvidia-smi
M3T2 The latency budget TTFT decomposition (queue + prefill + first token), TPOT/ITL, why p99 ≠ p50 × 2 Latency decomposer: instrument each stage of a request Mermaid: sequence diagram of request lifecycle A stacked-bar latency breakdown for one request
M3T3 The throughput ceiling Bandwidth-bound decode ceiling tok/s ≈ BW / bytes_per_param_read; where batching saturates Theoretical-vs-measured ceiling calculator Chart: theoretical ceiling vs measured, vs batch size The gap between theory and reality, explained
M3T4 $/1M tokens GPU-hour cost → cost per million input and output tokens at a given SLO; utilization as the hidden multiplier Cost model spreadsheet/notebook, parameterized Chart: $/1M tokens vs concurrency, with SLO cutoff Your own cost curve — the artifact you show leadership

Labs

Lab Goal Success criterion
LAB-M1 Profile your own rig Health report + measured bandwidth within 15% of a known-good reference
LAB-M2 Build a generation loop from scratch Generates coherent text with a working KV cache; matches HF generate() token-for-token at temperature 0
LAB-M3 Predict-then-measure Predict VRAM and throughput for 3 unseen configs; all three within 15% before you're allowed to proceed

LAB-M3 is the gate. It's the difference between having read this and knowing it.


Capstone

Deliverable: llm-sizing-tool — a small CLI + notebook that takes a Hugging Face model ID, a workload description (input/output length distribution, concurrency, SLO), and hardware spec, and outputs:

  • VRAM breakdown and whether it fits
  • Recommended TP degree and max batch size
  • Predicted TTFT, TPOT, and throughput
  • $/1M input and output tokens
  • The binding constraint, named ("you are KV-cache-capacity bound")

Rubric: predictions within 20% of measured reality on three models you haven't tested yet. Everything in Skill 01 will be used to falsify and refine this tool — that's the point.


Assessment

Tier Count Example
Recall 7 State the KV cache size formula and define every term
Apply 12 Given this config.json and 60 concurrent 8k-token users, does it fit on one A100 80GB? Show working
Design 4 Your model is KV-bound. Rank five interventions by expected impact and justify the ordering

Practical challenge: given an unfamiliar model config and a hardware spec, produce a complete sizing recommendation in 30 minutes, then verify it on real hardware.


Asset inventory

Type Count
Theory files 15
Code artifacts 15
Mermaid diagrams ~20
Generated images ~10
Charts ~12
Labs 3
Capstone 1

Primary sources

  • Karpathy — Let's build GPT + nanoGPT
  • vLLM PagedAttention paper (arXiv 2309.06180) — read §2 for the memory analysis
  • Efficient Memory Management for LLM Serving + the FlashAttention papers (for the SRAM/HBM argument)
  • GPU MODE lecture series (roofline, memory hierarchy)
  • Hugging Face transformers source: modeling_llama.py — read the actual attention implementation
  • NVIDIA A100 architecture whitepaper (for the hardware numbers you'll benchmark against)