Skill 00 — Foundations Assessment¶
Current scope: M1 — GPU & Systems Literacy
Complete the questions without looking up formulas. Show units and assumptions for every numeric answer.
Recall¶
- Name the four GPU-local memory tiers introduced in M1 and state the sharing scope of each.
- Define arithmetic intensity and the roofline ridge point.
- Distinguish model FLOPs utilization from hardware FLOPs utilization.
- Explain the range-versus-resolution trade between BF16 and FP16.
- Define all-gather, reduce-scatter, and all-reduce in terms of input and output ownership.
Apply¶
- A device sustains 1.8 TB/s and has a 270 TFLOP/s bf16 ceiling. Compute its ridge point.
- A copy reads and writes a 2 GiB buffer in 2.5 ms. Compute effective bandwidth in decimal GB/s.
- Estimate the HBM arithmetic intensity of a batch-one bf16 matrix-vector multiply when weight traffic dominates. How does the optimistic value change at batch 16?
- For a square bf16 GEMM, use the minimal-traffic model to estimate arithmetic intensity at
n=3072. Classify it against the ridge from question 6. - A W4A16 checkpoint is 3.5× smaller than BF16 but only 1.3× faster at batch one. Give three concrete sources of the missing speedup.
- For ring all-reduce with four ranks and a 400 MiB tensor, estimate the bytes moved across each rank’s links using the standard traffic model.
nvidia-smiworks,torch.__version__ends in+cpu, andtorch.version.cudaisNone. Which contract is broken, and why is installingnvccalone insufficient?
Design¶
- A decode service has low tensor-core utilization, high HBM bandwidth, and improves throughput almost linearly from batch 1 to 16. Choose the next two optimization experiments and defend them with a roofline argument.
- A two-GPU tensor-parallel model is slower than one GPU. Design the minimum experiment that distinguishes slow interconnect, startup-dominated collectives, and inefficient smaller GEMMs.
- Two nominally identical hosts differ by 18% in throughput. Design a health evidence bundle that can attribute the gap before any model code is changed.
Practical challenge¶
Complete LAB-M1 on real hardware. Submit the raw JSON, generated charts, health report, topology snapshot, exact commands, and a one-page explanation of the binding constraint. The result fails if it substitutes a CPU diagnostic or modeled NCCL curve for a GPU measurement.
Answer key
### Recall 1. Registers are per thread; shared memory/L1 is near and scoped to an SM/block according to the resource; L2 is device-wide; HBM is device capacity behind memory controllers. Exact cache behavior is architecture-dependent. 2. Arithmetic intensity is useful FLOPs divided by bytes crossing a named memory boundary. The ridge is `peak FLOP/s ÷ sustained byte/s`, where the sloped memory roof meets the flat compute roof. 3. MFU counts useful model operations relative to matching hardware peak. HFU counts executed hardware floating-point work and may include recomputation or other work that does not advance the model. 4. BF16 retains 8 exponent bits and has wider range with 7 fraction bits. FP16 has 5 exponent bits and 10 fraction bits, giving finer in-range resolution but earlier overflow. 5. All-gather turns one shard per rank into the full concatenation on every rank. Reduce-scatter reduces full contributions and leaves one reduced shard per rank. All-reduce leaves the full reduced tensor on every rank and is commonly decomposed into reduce-scatter plus all-gather. ### Apply 6. `270/1.8 = 150 FLOP/byte` after matching tera-units. 7. Traffic is `4 GiB = 4 × 2^30` bytes. Divide by `0.0025 s` and by `10^9`: approximately `1718 GB/s`. 8. About `1 FLOP/byte` at batch one: two FLOPs per bf16 weight and two bytes per weight. Ideal weight reuse raises it toward `16 FLOP/byte` at batch 16 before activation and other traffic. 9. Square GEMM intensity is approximately `n/3` for bf16 under the minimal model, so `3072/3 = 1024 FLOP/byte`. It lies to the compute side of a 150 FLOP/byte ridge. 10. Valid sources include scale/zero-point metadata, unpack/dequantization work, a kernel that fails to saturate bandwidth, unsupported shape fallback, activation traffic, launch overhead, and the operation moving toward a compute limit. 11. `2(N-1)/N × S = 1.5 × 400 MiB = 600 MiB` sent/received per rank under the ring traffic model. 12. The installed PyTorch build has no CUDA runtime support. `nvcc` is a toolkit compiler; it does not change how the existing PyTorch wheel was built. ### Design 13. Increase useful batch/reuse while checking the latency SLO, and test weight compression with a native kernel. Both move or raise the bandwidth-side limit; kernel fusion is another defensible experiment if profiler evidence shows intermediate HBM traffic. 14. Measure the layer’s local GEMMs at full and sharded shapes, then sweep isolated all-reduce message sizes while recording topology. The GEMM delta isolates shape efficiency; the curve’s knee isolates startup; its large-message plateau and forced-P2P control isolate fabric effects. 15. Capture stable GPU UUID/SKU/form factor, driver, framework CUDA build, topology, device mask, active processes, MIG/ECC state, clocks under load, power/temperature/throttle reasons, retired pages/errors, and the same memory-bandwidth control. Compare identical commands and workload inputs.M2 — Transformer Internals from the Inference Angle¶
Recall¶
- Name the two update branches in a decoder block and explain the residual stream.
- Write the KV-cache byte formula and define every term.
- Distinguish MHA, GQA, MQA, and MLA by the state they cache or share.
- Define prefill, decode, TTFT, and TPOT.
- Explain how RoPE makes a query–key score sensitive to relative position.
- Distinguish total parameters, active parameters, and expert load in an MoE.
Apply¶
- A model has 32 layers, 32 query heads, eight KV heads, head dimension 128, batch 20, context 8,192, and bf16 cache. Compute KV-cache size.
- For
d=4096,d_ff=14336, and a gated MLP, estimate MLP parameters per layer. - A model changes from 32 KV heads to four while all other cache terms stay fixed. What is the capacity ratio, and what evidence is still missing?
- A request has 4,000 prompt tokens and 20 output tokens. Name the phase responsible for initial delay and the phase repeated 20 times.
- Linear RoPE scaling maps a requested 16k range into a trained 4k range. What happens to a positional distance of 800?
- An eight-expert top-2 layer routes 35% of assignments to one expert. Why can average GPU utilization hide the resulting bottleneck?
Design¶
- Design a test that catches a KV cache whose shapes and bytes are correct but positions are wrong.
- You must serve long conversations at high concurrency. Compare GQA, KV quantization, shorter context, and more GPU memory; state the evidence needed before choosing.
- A 30B-total/3B-active MoE is slower than a dense 8B model. Build a measurement plan that can attribute the gap.
M2 practical challenge¶
Complete LAB-M2. Submit parameter predictions, token-level cache equivalence, cache byte validation, separate phase measurements, RoPE position tests, and—where used—expert-load evidence.