LAB-M1 — Profile Your Own Rig¶
Goal: Produce a reproducible GPU health and performance dossier that explains memory, compute, precision, and interconnect limits.
Time: 150 min Hardware: 1 CUDA GPU; 2 or 4 GPUs for the collective extension
Success criteria¶
- [ ]
health.jsonnames the driver, framework CUDA build, visible devices, topology, active processes, clocks, power, and ECC state—or records why a field is unavailable. - [ ] Three large-buffer bandwidth medians vary by less than 5% and the median is within 15% of a known-good result from the same GPU SKU and form factor.
- [ ] The roofline uses the measured bandwidth and matching dense-compute peak; every measured GEMM is classified before the run and explained afterward.
- [ ] The precision table includes three real checkpoint tensors and explains at least one range, outlier, or scaling reversal.
- [ ] On a multi-GPU host, both normal and peer-to-peer-disabled all-reduce curves identify a startup-to-bandwidth knee. On a one-GPU host, the limitation is recorded rather than simulated.
Setup¶
From 00-foundations/, synchronize the skill environment:
Record the exact GPU SKU and form factor. Obtain its published HBM bandwidth and dense compute peak for the dtype you will test. Obtain one known-good large-buffer result from the same hardware configuration; a different capacity or PCIe/SXM variant is not interchangeable.
Steps¶
-
Run
code/M1T5-toolchain-diagnostics/run.ps1. Resolve every failed contract that should pass on this host. Archive the report before load testing. -
Run M1T1 with
--size-mib 4, then report that number as though it were HBM bandwidth. This is the wrong path: the result is dominated by caching and launch behavior. Falsify it by sweeping 4, 64, 256, and 1024 MiB, then choose the plateau rather than the maximum. -
Repeat the largest safe size three times with
--spec-gbpsset to the exact published peak. Capture clocks, power, temperature, and competing processes during the middle run. If variance exceeds 5%, diagnose before continuing. -
Feed the sustained median into M1T2 with
--bandwidth-gbps. Supply the dense peak for the dtype actually used. Before running, derive the ridge and classify every square shape. Add three production-shaped GEMMs toroofline.py; include one shape whose dimensions are deliberately awkward for tensor-core tiling. -
Select three tensors from one checkpoint: embedding, attention projection, and MLP projection. Run M1T3 separately on each. Compare FP-like representation error and symmetric INT8/INT4 error. Do not convert MSE directly into a model-quality claim.
-
If two or four GPUs are available on Linux, capture
nvidia-smi topo -mand run M1T4. Locate the message size where each curve leaves the startup regime. Explain the normal-versus-control gap using topology evidence. If NCCL is unavailable, preserve the missing measurement as an explicit limitation. -
Run M1T5 again under load. Diff the idle and active JSON reports. Confirm that low idle clocks rise under work and that no unexpected process shares the device.
Stretch¶
- Replace the square-only roofline with projection shapes from a model you serve, separated into prefill-like and decode-like batch-token dimensions.
- Compare an unfused sequence of elementwise kernels with a fused implementation and account for the removed HBM traffic.
- On four GPUs, compare world sizes 2 and 4 and predict the ring traffic multiplier before measuring.
Debrief¶
The small-buffer bandwidth run should be the most seductive wrong answer in the lab: it is fast, repeatable, and irrelevant to streaming a model larger than cache. The large-buffer plateau is lower because it exposes the HBM boundary you intended to measure.
The roofline should explain why a point can be both compute-side by arithmetic intensity and far below peak: the bound removes one excuse but does not certify kernel efficiency. The precision experiment should show that range and resolution produce different rankings from storage width. On multiple GPUs, the collective curve should reveal two systems—startup-limited and bandwidth-limited—inside one API call.
Your final dossier is credible only if another engineer can reproduce it from the recorded hardware, software, commands, and raw JSON. A screenshot of nvidia-smi plus one throughput number does not pass.