Prompt
Why is my distributed training so slow between nodes?
Latest observation
Distributed training between nodes is usually slow because communication becomes the bottleneck, not the GPUs. The most common causes are:
1) Network bandwidth is too low
If your nodes communicate over 1/10 GbE instead of InfiniBand or 25/100+ GbE, gradient synchronization can dominate runtime.
2) Network latency is high
Even if bandwidth is decent, lots of small all-reduce / collective operations can be slowed a lot by latency.
3) Inefficient communication pattern
Framework settings may cause too many synchronizations, too-frequent gradient exchanges, or poor overlap between compute and communication.
4) Weak scaling issue
If you keep the per-GPU batch size too small, each node does little compute but still pays the same communication cost, so the communication/compute ratio gets worse.
5) CPU / dataloader bottleneck
Sometimes the GPUs are “slow between nodes” only because one node is waiting on data loading, preprocessing, or host-side synchronization.
6) PCIe / NCCL / topology problems
Communication may be slowed by:
- GPUs not using the fastest interconnect
- NCCL falling back to a slower path
- bad NUMA placement
- traffic going through the wrong NIC/PCIe root complex
- oversubscribed switches
7) Serialization / parameter size
Models with many small tensors can create overhead from many collective calls. Large models can be bandwidth-bound.
8) Mixed networking congestion
Other jobs, storage traffic, or VM/network contention can reduce effective throughput.
What to check first
- Measure GPU utilization
- If GPUs are often idle, you likely have communication or input pipeline bottlenecks.
- Check interconnect
- Confirm you’re using the intended network fabric and speed.
- Benchmark collectives
- Run NCCL tests or your framework’s communication benchmark.
- Increase batch size / gradient accumulation
- This often improves compute-to-communication ratio.
- Inspect NCCL logs
- Look for topology or transport fallbacks.
- Profile dataloading
- Make sure CPUs and storage aren’t the bottleneck.
Common fixes
- Use larger batch sizes or gradient accumulation
- Enable mixed precision
- Tune DDP / all-reduce bucket sizes
- Ensure NCCL is configured correctly
- Use faster networking (25/100GbE, InfiniBand)
- Improve data pipeline and pin CPUs/NUMA locality
- Reduce synchronization frequency where possible
If you want, I can help you diagnose it systematically. Tell me:
- framework (PyTorch / TensorFlow / JAX)
- number of nodes and GPUs per node
- network type/speed
- whether GPUs are idle or busy
- any NCCL or training logs