Prompt
I'm building a distributed training cluster—what should I look for in GPUs and networking?
Latest observation
For a distributed training cluster, the main goal is to keep all GPUs fed with data and synchronized efficiently. The two biggest bottlenecks are usually GPU interconnect and networking between nodes, but storage and CPU/RAM also matter.
1) GPU considerations
a) VRAM capacity
More VRAM lets you:
- train larger models
- use bigger batch sizes
- keep more activations / optimizer states in memory
For distributed training, VRAM still matters a lot even if you scale across nodes, because each GPU typically holds a shard or replica plus local activations.
b) GPU compute performance
Look at:
- Tensor Core performance for mixed precision training
- support for FP16, BF16, FP8 depending on your stack
- sustained throughput, not just peak specs
For modern deep learning, BF16 support is especially useful if available.
c) GPU-to-GPU interconnect
Within a node, fast GPU-to-GPU communication matters for:
- data parallel all-reduce
- tensor/pipeline parallelism
- sharded optimizers
Prefer systems with:
- NVLink / NVSwitch if available
- strong PCIe topology if not
This can matter as much as raw compute for distributed workloads.
d) Memory bandwidth
High bandwidth HBM helps with transformer training and other memory-heavy workloads.
e) Software ecosystem
Make sure the GPU line is well supported by:
- CUDA / ROCm
- PyTorch / TensorFlow / JAX
- NCCL or equivalent collectives
- mixed-precision and distributed libraries
A slightly slower GPU with excellent software support can outperform a “faster” one that’s harder to use.
f) Power and thermals
Distributed clusters are often limited by:
- rack power density
- cooling
- sustained clocks under load
Check:
- TDP
- form factor
- airflow/cooling requirements
- whether your datacenter can handle the density
2) Networking considerations
This is critical for multi-node training.
a) Bandwidth
Look for high node-to-node throughput:
- 100 GbE is a common baseline
- 200/400 GbE or InfiniBand is better for serious scaling
If gradients or activations must move frequently, bandwidth can become a bottleneck quickly.
b) Latency
Low latency matters for:
- synchronous training
- collectives like all-reduce/all-gather
- pipeline/tensor parallel communication
InfiniBand usually has an edge here, but modern Ethernet can be good if tuned properly.
c) Congestion behavior
At scale, network quality matters more than peak speed. Consider:
- packet loss handling
- congestion control
- oversubscription ratio
- topology design
A fast network that congests badly can perform worse than a slower but cleaner one.
d) RDMA support
If possible, use:
- InfiniBand
- or RoCEv2 on Ethernet
RDMA reduces CPU overhead and improves communication efficiency for training.
e) Topology
Good topology matters:
- leaf-spine is typical
- avoid oversubscription if you can
- keep nodes in the same high-bandwidth fabric domain when possible
If you do tensor parallelism or frequent cross-node communication, topology can significantly affect scaling efficiency.
f) NIC quality and placement
Check:
- number of NICs per node
- PCIe lane placement
- whether NICs share bandwidth with GPUs or storage
- driver maturity and RDMA support
A NIC plugged into a weak PCIe slot can bottleneck an otherwise strong node.
3) Don’t forget storage and CPU
Storage
Training jobs need fast access to datasets and checkpoints:
- local NVMe for staging/cache
- parallel filesystem or object store for shared data
- enough throughput to avoid data starvation
CPU and RAM
You need enough CPU to handle:
- dataloading
- preprocessing
- augmentation
- communication stack overhead
- orchestration
Insufficient CPU can starve GPUs even if the GPUs and network are excellent.
4) Practical buying priorities
If you want the short version, prioritize in this order:
- Enough GPU VRAM for your target model
- Fast interconnect inside the node (NVLink/NVSwitch if possible)
- High-bandwidth, low-latency networking between nodes
- InfiniBand or 200/400GbE if budget allows
- Good software support
- Adequate storage and CPU feeding the GPUs
- Power/cooling headroom
5) Questions to ask vendors
- What is the measured training scaling efficiency for multi-node workloads?
- Is the system optimized for data parallel, tensor parallel, or both?
- What’s the GPU-to-GPU topology within the node?
- What network fabric is used, and what is the oversubscription ratio?
- Do you support RDMA and which collectives stack?
- How much of peak performance is sustained under real training loads?
- What are the power and cooling requirements at full load?
6) Rule of thumb
- If you’re training mostly single-node: GPU VRAM + NVLink matter most.
- If you’re training multi-node: network quality can make or break scaling.
- If you’re training large language models or heavy distributed workloads: choose the fastest fabric you can justify, because communication becomes a major fraction of runtime.
If you want, I can also help you design a cluster for a specific workload, like:
- LLM training
- vision models
- recommendation systems
- 8-node / 16-node / 64-node clusters
- budget-constrained vs performance-first setups