Prompt
Why is my distributed training hanging on GPU startup?
Latest observation
A distributed training job that “hangs on GPU startup” is usually stuck during process/GPU initialization, not actual model training. Common causes are:
Most common reasons
1) NCCL / collective communication init is blocked
If you’re using PyTorch DDP, Horovod, DeepSpeed, JAX, etc., the job may hang while initializing the process group or performing the first all-reduce.
Typical causes:
- wrong
world_size,rank,local_rank - one or more ranks failed before init, so others wait forever
- networking/firewall issues between nodes
- NCCL can’t find a working interface
2) GPU visibility / device assignment mismatch
Processes may be waiting because:
- multiple ranks are trying to use the same GPU
CUDA_VISIBLE_DEVICESis inconsistent across nodeslocal_ranknot mapped correctly to GPU index- too many processes launched for available GPUs
3) Driver / CUDA / NCCL mismatch
A mismatch between:
- NVIDIA driver
- CUDA runtime/toolkit
- PyTorch build
- NCCL version
can cause initialization to stall or error very late.
4) Forking after CUDA initialization
If you use Python multiprocessing and initialize CUDA before forking, child processes can hang.
Common pattern:
- use
spawninstead offork - ensure CUDA is not touched in the parent before worker launch
5) One rank is doing something slow or crashed
If rank 0 is loading data, downloading checkpoints, or stuck on CPU work while others hit barrier/init, it looks like a hang.
6) DataLoader worker startup deadlock
Sometimes the “GPU startup” is actually the first batch pipeline:
- too many workers
- bad
collate_fn - dataset code deadlocks
pin_memory/shared memory issues
7) NCCL/network transport issues on multi-node jobs
On multi-node clusters:
- bad
MASTER_ADDR/MASTER_PORT - blocked port
- incorrect interface selection
- IB/RDMA misconfiguration
Quick checks
Single-node
- Confirm each process gets a unique GPU.
- Print:
- rank
- local rank
CUDA_VISIBLE_DEVICEStorch.cuda.device_count()
- Try
nproc_per_node == number_of_GPUs
Multi-node
- Verify all nodes can reach
MASTER_ADDR:MASTER_PORT - Ensure identical software stack on all nodes
- Check clocks/hostnames/DNS if using hostnames
Useful debugging env vars
For PyTorch/NCCL:
export NCCL_DEBUG=INFO
export NCCL_DEBUG_SUBSYS=INIT,ENV
export TORCH_DISTRIBUTED_DEBUG=DETAIL
Sometimes also useful:
export NCCL_IB_DISABLE=1
export NCCL_P2P_DISABLE=1
export NCCL_SOCKET_IFNAME=eth0 # or your actual interface
If disabling IB/P2P makes it work, the issue is likely transport/topology-related.
Practical debugging steps
- Run with 1 GPU, 1 process.
- Then try 1 node, all GPUs.
- Then scale to multi-node.
- Add logging before and after:
- process launch
init_process_groupcuda.set_device- first forward pass
- first backward/all-reduce
- Test a minimal NCCL script separate from your training code.
What I’d need to pinpoint it
If you want, paste:
- framework (PyTorch/DeepSpeed/Horovod/etc.)
- single-node or multi-node
- launch command
- relevant logs around startup
- GPU count and
CUDA_VISIBLE_DEVICES - whether it hangs before or after
init_process_group
And I can help narrow it down quickly.