Prompt
Building an on-prem AI server rack with NVIDIA GPUs
Latest observation
Building an on-prem AI server rack with NVIDIA GPUs is very doable, but the “right” design depends heavily on your target workload, budget, power/cooling limits, and whether you’re optimizing for training, fine-tuning, or inference.
Here’s a practical framework to design it.
1) Start with the workload
Decide what you’re building for:
- Inference only: lower power, simpler networking, less memory pressure
- Fine-tuning / LoRA / DPO: needs solid GPU VRAM and good storage throughput
- Full training: highest demands on GPU interconnect, CPU, RAM, storage, cooling
Key questions:
- What model sizes? (7B, 13B, 70B, multimodal, vision, etc.)
- How many concurrent users/requests?
- Latency target?
- Batch vs real-time?
- Single-node or multi-node cluster?
2) Pick the GPU class
For NVIDIA, your options usually fall into these buckets:
Best value / prosumer
- RTX 4090 / 5090-class depending on availability
- Great for inference and light fine-tuning
- Limitations:
- No NVLink in most consumer cards
- Lower VRAM per dollar than datacenter cards
- Cooling/power can be awkward in dense racks
Datacenter / workstation
- L40S
- A100
- H100 / H200
- Better suited for rack deployment, ECC memory, better thermals, more consistent uptime
- Much better if you need multi-GPU scaling or 24/7 operation
Rule of thumb
- If you want maximum price/performance for a small rack: consumer GPUs can work
- If you want reliability, dense deployment, and easier operations: datacenter GPUs are worth it
3) Decide node architecture
A clean approach is usually:
Option A: 1–2 big GPU nodes
Good for:
- Small teams
- Simple ops
- Shared storage and modest scale
Typical node:
- 4x GPUs
- Dual EPYC or Xeon
- 256–1024 GB RAM
- NVMe storage
- 25/50/100GbE networking
Option B: several smaller GPU nodes
Good for:
- Horizontal scaling
- Fault isolation
- Easier maintenance
- Mixed workloads
Typical:
- 1–2 GPUs per node
- Central orchestration with Kubernetes, Slurm, or simple Docker/Compose
4) CPU, RAM, and motherboard considerations
For AI servers, GPU is the headline, but the rest matters a lot.
CPU
- Prefer AMD EPYC for many PCIe lanes and strong server support
- Intel Xeon also works, but EPYC is often the easier fit for multi-GPU
RAM
- Minimum: 256 GB
- Better: 512 GB+ for data loading, caching, preprocessing, and large jobs
Motherboard
Look for:
- Enough PCIe x16 physical slots
- Proper lane bifurcation support
- Confirmed GPU compatibility
- IPMI/BMC remote management
- Good spacing for air-cooled cards
5) Power planning
This is often the biggest surprise.
Estimate carefully
Example rough GPU power:
- RTX 4090: ~450W each
- L40S: ~350W each
- H100 SXM systems can be far higher, especially in dense configs
A 4-GPU node can easily draw:
- 2.0 kW to 3.5 kW+ under load
Add:
- CPU(s)
- Memory
- NVMe
- Fans
- PSU inefficiency
Recommendations
- Use redundant high-efficiency PSUs
- Plan for 208V or 240V power
- Check rack circuit capacity early
- Don’t assume a standard office circuit will work
6) Cooling and rack design
High-density AI racks are as much about heat as compute.
Air cooling
Works if:
- You have good aisle containment
- Sufficient rack airflow
- GPUs are not overly cramped
Liquid cooling
Worth considering if:
- You’re running high-power datacenter GPUs
- Rack density is high
- Noise or thermal constraints are tight
Practical advice
- Leave space for airflow between nodes
- Use blanking panels in unused U slots
- Monitor inlet temps, not just ambient room temp
7) Storage
AI workloads often bottleneck on data, not compute.
Local scratch
- Fast NVMe SSDs
- RAID 0/10 depending on risk tolerance
Shared storage
- NAS/SAN/NVMe-oF if you have multiple nodes
- Ceph or Lustre for larger environments
- For smaller setups, a good 10/25/100GbE NAS may be enough
Best practice
- Keep datasets on fast shared storage
- Keep training cache and checkpoints on local NVMe
8) Networking
Networking matters more once you have more than one node.
- 25GbE: decent starting point
- 50/100GbE: better for serious multi-node training
- InfiniBand: ideal for low-latency distributed training if budget allows
For inference clusters:
- 10/25GbE is often sufficient unless traffic is heavy
9) Software stack
Common stack choices:
- OS: Ubuntu Server or Rocky Linux
- Drivers/CUDA: match carefully with framework versions
- Container runtime: Docker + NVIDIA Container Toolkit
- Orchestration:
- Kubernetes for service-heavy environments
- Slurm for HPC/training jobs
- Plain systemd/Docker for smaller deployments
Frameworks:
- PyTorch
- TensorFlow
- vLLM / TGI / TensorRT-LLM for inference
- NCCL for multi-GPU communication
10) Example build profiles
Budget 1-node inference box
- 2x RTX 4090
- 1x EPYC CPU
- 256 GB RAM
- 2–4 TB NVMe
- 10/25GbE
- ~1.2–1.5 kW load
Mid-range fine-tuning node
- 4x L40S or 4x 4090-class
- Dual EPYC
- 512 GB RAM
- Multiple NVMe drives
- 25/100GbE
- ~2–3 kW load
Higher-end training node
- 4–8x H100/H200-class GPUs
- Dual-socket server platform
- 512 GB to 1 TB RAM
- 100/200GbE or InfiniBand
- Liquid cooling often preferred
- Significant power/cooling infrastructure required
11) Common mistakes to avoid
- Underestimating power draw
- Buying GPUs that don’t physically fit together
- Ignoring cooling and rack airflow
- Skipping ECC/remote management
- Not planning for spare parts
- Mixing consumer GPUs into a dense rack without checking thermals
- Buying storage/networking too weak for the compute
12) A good first-step plan
If you’re just starting, a sensible path is:
- Define workload and budget
- Decide on 1-node vs multi-node
- Choose GPU class
- Size power/cooling first
- Choose server chassis and motherboard
- Add RAM/NVMe/networking
- Validate software stack on a pilot node
- Expand after real workload testing
If you want, I can help you with one of these next:
- a specific bill of materials
- a rack power/cooling plan
- a GPU comparison for your budget
- a single-node design or multi-node cluster design