Prompt
Why do my distributed GPU jobs keep timing out during NCCL setup?
Latest observation
Oct 11, 2026
OpenAI APIWeb search: off
Distributed GPU jobs most often time out during NCCL setup because the processes can’t successfully form a low-latency communication group fast enough. The usual causes fall into a few buckets:
Common reasons
-
Network/interface mismatch
- NCCL may pick the wrong network interface.
- On multi-NIC machines, it might choose a management or docker interface instead of the high-speed fabric.
-
Firewall / security group / port blocking
- Rank-to-rank connections may be blocked.
- Even if SSH works, NCCL may need additional ports or unrestricted intra-node/inter-node traffic.
-
Mismatched node configuration
- Different CUDA/NCCL/driver versions across nodes.
- Different GPU counts or topology assumptions.
- One node is slower or missing GPUs.
-
Launcher / rendezvous issues
torchrun, MPI, SLURM, or your rendezvous backend may be misconfigured.- Wrong
MASTER_ADDR,MASTER_PORT,WORLD_SIZE,RANK, or timeout settings.
-
Slow initialization / overloaded system
- Startup takes too long due to heavy CPU load, GPU contention, or filesystem delays.
- NCCL init can stall if a process is delayed and others are waiting.
-
IB/RDMA problems
- InfiniBand / RoCE / RDMA misconfiguration can cause hangs or retries.
- NCCL may attempt IB and fail slowly instead of falling back cleanly.
-
Container / namespace issues
- Docker/Kubernetes network isolation can hide the correct interfaces or block peer access.
- NCCL sees different hostnames/IPs than expected.
-
Topology or GPU visibility problems
CUDA_VISIBLE_DEVICESdiffers across ranks.- Incorrect mapping of local rank to device can make one process hang.
What to check first
- Confirm all ranks can reach each other on the rendezvous port.
- Set NCCL debug logs:
export NCCL_DEBUG=INFO export NCCL_DEBUG_SUBSYS=INIT,NET - Verify the interface NCCL uses:
export NCCL_SOCKET_IFNAME=eth0 # or your real high-speed NIC - If IB is suspect, test disabling it:
export NCCL_IB_DISABLE=1 - Make sure
MASTER_ADDR,MASTER_PORT,WORLD_SIZE, andRANKare correct. - Check that all nodes have consistent driver/CUDA/NCCL versions.
Practical debugging approach
- Run a minimal NCCL test (
nccl-testslikeall_reduce_perf). - Try one node vs multiple nodes to isolate whether it’s networking or local GPU setup.
- Disable IB temporarily to see if the timeout disappears.
- Increase initialization timeout if startup is just slow.
- Inspect logs from every rank, not just rank 0.
If you want, I can help you debug it systematically if you paste:
- your launch command,
- the NCCL/torch distributed error,
- node count, GPU count, and network type (Ethernet/IB/RoCE),
- and whether it fails only multi-node or also single-node.