Prompt
Building a multi-node training setup with InfiniBand
Latest observation
Oct 11, 2026
OpenAI APIWeb search: off
Here’s a practical blueprint for building a multi-node training setup with InfiniBand (IB), focused on distributed deep learning.
1) What you need
- Servers with IB-capable NICs: Typically Mellanox/NVIDIA ConnectX cards.
- IB switch: Or direct connect for very small setups, but a switch is standard.
- Compatible cables/transceivers: QSFP/QSFP28/QSFP56 depending on generation.
- GPU nodes: If training on GPUs, ensure each node has enough PCIe lanes / NVLink where applicable.
2) Network topology
For training, keep it simple:
- Put all training nodes on the same IB fabric.
- Use a single high-speed fabric for:
- data parallel collectives (all-reduce)
- parameter synchronization
- optionally storage access if using RDMA-capable storage
If you also need Ethernet:
- Use Ethernet for management, SSH, monitoring, and job orchestration.
- Use IB only for training traffic.
3) Software stack
Typical stack:
- OS: Ubuntu/Rocky/CentOS supported by your vendor.
- OFED / RDMA drivers:
- NVIDIA/Mellanox OFED or inbox RDMA drivers
- MPI if your framework uses it:
- OpenMPI, MPICH, or NCCL-based setup
- Deep learning framework:
- PyTorch Distributed
- TensorFlow MultiWorkerMirroredStrategy
- Horovod
- NCCL for GPU collectives on NVIDIA GPUs
4) Basic setup steps
A. Verify hardware
On each node:
lspci | grep -i mell
ibv_devinfo
ibstat
B. Install RDMA stack
Install the vendor OFED package or the distro RDMA stack, then load drivers:
lsmod | grep mlx
systemctl status opensmd
(Depending on your environment, the subnet manager may run on the switch or a host.)
C. Configure IP over IB if needed
If you want regular IP networking over IB:
ip link show
ip addr add 10.10.0.1/24 dev ib0
ip link set ib0 up
But for training performance, many stacks use RDMA directly, not IPoIB.
D. Test connectivity and bandwidth
Use IB tools:
ibping
ibv_rc_pingpong
perftest
Examples:
ib_write_bw
ib_read_bw
5) Deep learning configuration
PyTorch distributed
Common approach with NCCL:
- Set node ranks and master address
- Use all GPUs per node
- Ensure NCCL sees IB devices
Example env vars:
export NCCL_DEBUG=INFO
export NCCL_IB_DISABLE=0
export NCCL_SOCKET_IFNAME=eth0
export NCCL_IB_HCA=mlx5_0
Launch with torchrun:
torchrun \
--nnodes=2 \
--nproc_per_node=8 \
--node_rank=0 \
--master_addr=10.0.0.1 \
--master_port=29500 \
train.py
Horovod
mpirun -np 16 \
-H node1:8,node2:8 \
-bind-to none -map-by slot \
-x NCCL_DEBUG=INFO \
-x NCCL_IB_DISABLE=0 \
python train.py
6) Performance tuning
- Use GPUDirect RDMA if supported to avoid extra copies.
- Make sure PCIe topology is good:
- GPU and IB NIC should ideally be on the same PCIe root complex or well-optimized path.
- Pin CPU cores and memory when needed.
- Ensure the IB fabric runs at intended speed:
- 100G/200G/400G depending on cards/switches.
- Tune MTU if appropriate:
- Often 4092 or 2044 on IB fabrics, depending on environment.
- Check NUMA affinity:
- Bind processes to the CPU socket closest to the NIC/GPU.
7) Troubleshooting checklist
- Link not up: cables, optics, firmware, switch port config
- No RDMA devices: driver/OFED mismatch
- Poor performance: wrong NUMA binding, NCCL not using IB, PCIe bottleneck
- Collectives hang: firewall, rank misconfig, subnet manager issue, mixed library versions
- Mismatched MTU/speed: switch and NIC settings not aligned
Useful commands:
ibstat
iblinkinfo
rdma link show
nvidia-smi topo -m
nvidia-smi nvlink -s
8) Recommended “known good” baseline
For a first deployment:
- Same model server on all nodes
- Same GPU type per node
- Same IB NIC model/firmware
- Same OS and driver versions
- One subnet manager
- Use NCCL + PyTorch Distributed
- Keep management traffic on Ethernet, training traffic on IB
9) Minimal architecture
- 2–N nodes
- Each node:
- 8 GPUs
- 1–2 IB NICs
- 1 Ethernet management port
- IB switch
- Shared storage via Ethernet/NFS or RDMA-enabled storage
If you want, I can also provide:
- a step-by-step install guide for Ubuntu/RHEL,
- a PyTorch + NCCL multi-node example, or
- a hardware BOM/recommendation for a specific cluster size.