Prompt
How do I debug slow GPU-to-GPU communication in training?
Latest observation
To debug slow GPU-to-GPU communication during training, you want to determine whether the bottleneck is communication itself, synchronization, overlap, or topology/transport. Here’s a practical checklist.
1) First identify the communication pattern
Ask:
- Is this DDP/all-reduce heavy?
Common in data parallel training, especially with many small gradient buckets. - Is it point-to-point send/recv?
More common in pipeline or tensor parallelism. - Is communication happening across nodes or only within a node?
This changes whether PCIe/NVLink/NIC is the issue.
2) Measure where time is going
Use profiling before changing anything.
Useful tools
- PyTorch Profiler with CUDA/NCCL activities
- Nsight Systems (
nsys) for timeline analysis - NCCL debug logs:
NCCL_DEBUG=INFO NCCL_DEBUG_SUBSYS=INIT,GRAPH,COLL,NET - For PyTorch distributed, also inspect:
torch.cuda.synchronize()placementtorch.distributed.barrier()usage- dataloader stalls that may make comm look slow
What to look for
- Long gaps before/after NCCL calls
- Communication not overlapping with compute
- Many small collectives instead of fewer large ones
- Unexpected CPU-side delays before launching GPU ops
3) Check if you are actually synchronized too often
A very common mistake is forcing syncs accidentally:
.item()print(tensor)torch.cuda.synchronize()- frequent logging of GPU tensors
- debugging assertions that read tensors back to CPU
These can make GPU-GPU comm appear slow when the real issue is CPU-GPU synchronization.
4) Verify NCCL/topology setup
If using NVIDIA GPUs, NCCL is usually the backend. Slow comm often comes from bad topology or disabled fast paths.
Check:
- Are GPUs connected by NVLink or only PCIe?
- Are ranks mapped to GPUs on the same NUMA/socket correctly?
- Are you crossing CPU sockets unnecessarily?
- Is NCCL using the expected transport?
Helpful env vars
NCCL_DEBUG=INFO
NCCL_DEBUG_SUBSYS=INIT,GRAPH,NET
NCCL_P2P_DISABLE=0
NCCL_IB_DISABLE=0 # if using InfiniBand
Things that can hurt performance
- Using GPUs across different PCIe root complexes
- Bad process-to-GPU affinity
- Disabled P2P due to container/driver settings
- NCCL falling back to socket transport over IB/NVLink
5) Test the raw bandwidth/latency separately
Before debugging the training loop, benchmark the interconnect.
NCCL tests
Run nccl-tests:
all_reduce_perfsendrecv_perf
This helps distinguish:
- network/interconnect problem vs
- training code issue
If the benchmark is also slow, the issue is likely hardware, topology, or NCCL config.
6) Check message size and bucketization
For DDP, small gradient buckets can create lots of latency overhead.
Try:
- Increasing DDP bucket size
- Fusing smaller tensors
- Ensuring gradients are contiguous where possible
In PyTorch DDP:
DistributedDataParallel(model, bucket_cap_mb=...)
Too-small buckets = many launches and poor bandwidth utilization.
7) Determine whether overlap is happening
Good training overlaps compute with communication. If not, comm can dominate.
Symptoms of poor overlap
- Backward finishes, then comm starts
- Communication kernels show up serialized
- GPU utilization drops during comm phases
Potential causes
- Graph structure prevents early gradient readiness
- Requiring gradients too late
- Custom autograd ops delaying backward
- Calling
no_sync()incorrectly or not at all
8) Inspect CPU-side bottlenecks
Sometimes “GPU-GPU comm” is slow because the CPU orchestrating it is slow.
Check:
- One rank pinned to a busy CPU core
- Python overhead in the training loop
- Thread contention
- Dataloader starving the GPU
Use:
htop,mpstattorch.profilerCPU activity- pinning/affinity settings
9) Validate environment settings
A few knobs can accidentally degrade performance:
Common suspects
CUDA_VISIBLE_DEVICESmapping mismatchOMP_NUM_THREADStoo high- NCCL threading settings
- Container limits on shared memory or IPC
- MIG mode affecting peer access
Also check
- Driver/CUDA/NCCL version compatibility
- Whether peer access is enabled
- Whether GPUDirect RDMA is available/working for multi-node
10) Compare single-node vs multi-node
This helps narrow the issue:
- Fast on one node, slow across nodes → network/NIC/IB config
- Slow even on one node → PCIe/NVLink, process placement, bucketization, syncs
11) Use a minimal repro
Strip training down to:
- model forward/backward
- one all-reduce or one DDP iteration
- fixed tensor sizes
If the minimal case is fast, the problem is in your training loop.
If it’s slow, the problem is infrastructure/configuration.
Practical debugging sequence
If you want the fastest path:
- Profile with Nsight Systems or PyTorch Profiler
- Check for accidental syncs
- Run
nccl-tests - Enable NCCL debug logs
- Verify GPU affinity/topology
- Tune bucket sizes / reduce small collectives
- Test single-node vs multi-node
If you want, I can help you debug a specific setup
If you paste:
- your framework (PyTorch / JAX / DeepSpeed / etc.)
- single-node or multi-node
- number/type of GPUs
- NCCL logs or profiler screenshot
- your training/distributed setup code
…I can help narrow down the cause quickly.