Skip to content

Skill 00 — Foundations Assessment

Current scope: M1 — GPU & Systems Literacy

Complete the questions without looking up formulas. Show units and assumptions for every numeric answer.

Recall

  1. Name the four GPU-local memory tiers introduced in M1 and state the sharing scope of each.
  2. Define arithmetic intensity and the roofline ridge point.
  3. Distinguish model FLOPs utilization from hardware FLOPs utilization.
  4. Explain the range-versus-resolution trade between BF16 and FP16.
  5. Define all-gather, reduce-scatter, and all-reduce in terms of input and output ownership.

Apply

  1. A device sustains 1.8 TB/s and has a 270 TFLOP/s bf16 ceiling. Compute its ridge point.
  2. A copy reads and writes a 2 GiB buffer in 2.5 ms. Compute effective bandwidth in decimal GB/s.
  3. Estimate the HBM arithmetic intensity of a batch-one bf16 matrix-vector multiply when weight traffic dominates. How does the optimistic value change at batch 16?
  4. For a square bf16 GEMM, use the minimal-traffic model to estimate arithmetic intensity at n=3072. Classify it against the ridge from question 6.
  5. A W4A16 checkpoint is 3.5× smaller than BF16 but only 1.3× faster at batch one. Give three concrete sources of the missing speedup.
  6. For ring all-reduce with four ranks and a 400 MiB tensor, estimate the bytes moved across each rank’s links using the standard traffic model.
  7. nvidia-smi works, torch.__version__ ends in +cpu, and torch.version.cuda is None. Which contract is broken, and why is installing nvcc alone insufficient?

Design

  1. A decode service has low tensor-core utilization, high HBM bandwidth, and improves throughput almost linearly from batch 1 to 16. Choose the next two optimization experiments and defend them with a roofline argument.
  2. A two-GPU tensor-parallel model is slower than one GPU. Design the minimum experiment that distinguishes slow interconnect, startup-dominated collectives, and inefficient smaller GEMMs.
  3. Two nominally identical hosts differ by 18% in throughput. Design a health evidence bundle that can attribute the gap before any model code is changed.

Practical challenge

Complete LAB-M1 on real hardware. Submit the raw JSON, generated charts, health report, topology snapshot, exact commands, and a one-page explanation of the binding constraint. The result fails if it substitutes a CPU diagnostic or modeled NCCL curve for a GPU measurement.

Answer key ### Recall 1. Registers are per thread; shared memory/L1 is near and scoped to an SM/block according to the resource; L2 is device-wide; HBM is device capacity behind memory controllers. Exact cache behavior is architecture-dependent. 2. Arithmetic intensity is useful FLOPs divided by bytes crossing a named memory boundary. The ridge is `peak FLOP/s ÷ sustained byte/s`, where the sloped memory roof meets the flat compute roof. 3. MFU counts useful model operations relative to matching hardware peak. HFU counts executed hardware floating-point work and may include recomputation or other work that does not advance the model. 4. BF16 retains 8 exponent bits and has wider range with 7 fraction bits. FP16 has 5 exponent bits and 10 fraction bits, giving finer in-range resolution but earlier overflow. 5. All-gather turns one shard per rank into the full concatenation on every rank. Reduce-scatter reduces full contributions and leaves one reduced shard per rank. All-reduce leaves the full reduced tensor on every rank and is commonly decomposed into reduce-scatter plus all-gather. ### Apply 6. `270/1.8 = 150 FLOP/byte` after matching tera-units. 7. Traffic is `4 GiB = 4 × 2^30` bytes. Divide by `0.0025 s` and by `10^9`: approximately `1718 GB/s`. 8. About `1 FLOP/byte` at batch one: two FLOPs per bf16 weight and two bytes per weight. Ideal weight reuse raises it toward `16 FLOP/byte` at batch 16 before activation and other traffic. 9. Square GEMM intensity is approximately `n/3` for bf16 under the minimal model, so `3072/3 = 1024 FLOP/byte`. It lies to the compute side of a 150 FLOP/byte ridge. 10. Valid sources include scale/zero-point metadata, unpack/dequantization work, a kernel that fails to saturate bandwidth, unsupported shape fallback, activation traffic, launch overhead, and the operation moving toward a compute limit. 11. `2(N-1)/N × S = 1.5 × 400 MiB = 600 MiB` sent/received per rank under the ring traffic model. 12. The installed PyTorch build has no CUDA runtime support. `nvcc` is a toolkit compiler; it does not change how the existing PyTorch wheel was built. ### Design 13. Increase useful batch/reuse while checking the latency SLO, and test weight compression with a native kernel. Both move or raise the bandwidth-side limit; kernel fusion is another defensible experiment if profiler evidence shows intermediate HBM traffic. 14. Measure the layer’s local GEMMs at full and sharded shapes, then sweep isolated all-reduce message sizes while recording topology. The GEMM delta isolates shape efficiency; the curve’s knee isolates startup; its large-message plateau and forced-P2P control isolate fabric effects. 15. Capture stable GPU UUID/SKU/form factor, driver, framework CUDA build, topology, device mask, active processes, MIG/ECC state, clocks under load, power/temperature/throttle reasons, retired pages/errors, and the same memory-bandwidth control. Compare identical commands and workload inputs.

M2 — Transformer Internals from the Inference Angle

Recall

  1. Name the two update branches in a decoder block and explain the residual stream.
  2. Write the KV-cache byte formula and define every term.
  3. Distinguish MHA, GQA, MQA, and MLA by the state they cache or share.
  4. Define prefill, decode, TTFT, and TPOT.
  5. Explain how RoPE makes a query–key score sensitive to relative position.
  6. Distinguish total parameters, active parameters, and expert load in an MoE.

Apply

  1. A model has 32 layers, 32 query heads, eight KV heads, head dimension 128, batch 20, context 8,192, and bf16 cache. Compute KV-cache size.
  2. For d=4096, d_ff=14336, and a gated MLP, estimate MLP parameters per layer.
  3. A model changes from 32 KV heads to four while all other cache terms stay fixed. What is the capacity ratio, and what evidence is still missing?
  4. A request has 4,000 prompt tokens and 20 output tokens. Name the phase responsible for initial delay and the phase repeated 20 times.
  5. Linear RoPE scaling maps a requested 16k range into a trained 4k range. What happens to a positional distance of 800?
  6. An eight-expert top-2 layer routes 35% of assignments to one expert. Why can average GPU utilization hide the resulting bottleneck?

Design

  1. Design a test that catches a KV cache whose shapes and bytes are correct but positions are wrong.
  2. You must serve long conversations at high concurrency. Compare GQA, KV quantization, shorter context, and more GPU memory; state the evidence needed before choosing.
  3. A 30B-total/3B-active MoE is slower than a dense 8B model. Build a measurement plan that can attribute the gap.

M2 practical challenge

Complete LAB-M2. Submit parameter predictions, token-level cache equivalence, cache byte validation, separate phase measurements, RoPE position tests, and—where used—expert-load evidence.

M2 answer key 16. Attention mixes token information and the MLP transforms each token; both write updates into the continuing residual stream. 17. `2 × L × B × T × Hkv × dh × s`; the factor two is K and V, followed by layers, sequences, tokens, KV heads, elements per head, and bytes per element. 18. MHA has one KV head per query head; GQA shares each KV head among a query group; MQA shares one KV head among all queries; MLA caches a compressed latent representation according to its architecture. 19. Prefill processes the prompt and initializes cache; decode emits subsequent tokens. TTFT measures delay to the first token; TPOT measures cadence between output tokens. 20. Applying rotations by absolute positions makes the dot product contain the difference of rotation angles, hence relative displacement. 21. Total parameters occupy storage, active parameters participate for one token, and expert load counts routed token assignments that determine balance. 22. `2×32×20×8192×8×128×2 = 21,474,836,480 bytes = 20 GiB`. 23. `3×4096×14336 = 176,160,768` weights, excluding biases. 24. Cache becomes `4/32 = 1/8` of the MHA size. Missing evidence includes trained-weight compatibility, task quality, engine/kernel support, and end-to-end performance. 25. Prefill drives initial prompt work; decode repeats once for each of the 20 output steps. 26. Factor four compression maps distance 800 to scaled distance 200. 27. The busiest expert/rank determines synchronization time while aggregate utilization averages hot and waiting devices. 28. Compare cached output with full causal output after every token and shift absolute positions while preserving relative offsets; the first divergent token localizes the position error. 29. Derive capacity changes, then measure task quality, cache bandwidth, supported kernels, concurrency, TTFT/TPOT, and cost under the same workload. 30. Separate expert GEMMs, routing/grouping, weight reads, all-to-all dispatch/return, imbalance, batch composition, and dense baseline quality/shape.