Skip to content

M1T4 — Measure collective communication

On a Linux CUDA host with 2 or 4 visible GPUs:

WORLD_SIZE=2 ./run.sh

The runner measures the normal NCCL path and a second path with GPU peer-to-peer disabled, then writes two JSON files and results/all-reduce-bandwidth.png. It never fills in a missing curve with modeled data.

Success means identifying the startup-dominated knee and explaining the large-message bandwidth gap between the two modes. Before interpreting it, capture nvidia-smi topo -m; “forced PCIe” is an experiment control, not a claim that every byte followed an identical physical route.


View the runnable source on GitHub