M1T4 — Measure collective communication¶
On a Linux CUDA host with 2 or 4 visible GPUs:
The runner measures the normal NCCL path and a second path with GPU peer-to-peer disabled, then writes two JSON files and results/all-reduce-bandwidth.png. It never fills in a missing curve with modeled data.
Success means identifying the startup-dominated knee and explaining the large-message bandwidth gap between the two modes. Before interpreting it, capture nvidia-smi topo -m; “forced PCIe” is an experiment control, not a claim that every byte followed an identical physical route.