Prompt
How do I set up a GPU cluster for both training and inference?
Latest observation
Setting up a GPU cluster for both training and inference is mostly about balancing throughput, latency, scheduling, storage, and networking. A good design separates the two workloads while keeping them easy to manage.
1) Start with the workload split
Decide what you need for each:
-
Training
- Long-running jobs
- High GPU utilization
- Often needs multi-GPU or multi-node scaling
- Heavy storage and checkpoint I/O
-
Inference
- Low latency or high throughput
- Smaller, more frequent requests
- Benefits from autoscaling and serving optimizations
- Often uses different GPU types than training
If possible, dedicate some nodes to training and some to inference. Mixing them on the same nodes can cause noisy-neighbor issues.
2) Choose the cluster architecture
Common patterns
A. Separate pools in one cluster
- One Kubernetes cluster or one Slurm cluster
- Distinct node groups:
- training nodes
- inference nodes
- Easier shared management
- Good if your team is small
B. Separate clusters
- One cluster for training, one for inference
- Best isolation
- More operational overhead
- Good for strict SLOs or different security needs
For most teams, one cluster with separate node pools is the best start.
3) Pick hardware
GPU selection
- Training: prioritize VRAM, interconnect, multi-GPU performance
- Examples: NVIDIA H100, A100, L40S depending on budget/workload
- Inference: prioritize cost/performance and latency
- Examples: L4, L40S, A10, sometimes H100 if needed
Other hardware
- CPU: enough cores to feed GPUs
- RAM: especially important for data loading and inference preprocessing
- Storage: local NVMe on workers is very helpful
- Networking: at least 25/50/100 GbE; InfiniBand or RoCE for distributed training
4) Use a scheduler/orchestrator
You need a system to allocate GPUs cleanly.
Options
Kubernetes
- Good for mixed training + inference
- Strong ecosystem for serving, autoscaling, observability
- Use:
- NVIDIA GPU Operator
- node pools/taints/tolerations
- namespaces and quotas
- Horizontal Pod Autoscaler / Cluster Autoscaler
Slurm
- Excellent for HPC-style training
- Great for batch and multi-node jobs
- Less convenient for production inference
Hybrid
- Kubernetes for inference, Slurm for training
- Common in larger orgs
If you want both on one platform, Kubernetes is usually the most flexible.
5) Set up GPU scheduling
On Kubernetes, configure:
- NVIDIA drivers on nodes
- NVIDIA device plugin or GPU Operator
- Node labels like:
workload=trainingworkload=inference
- Taints/tolerations so inference pods don’t land on training nodes and vice versa
- Resource requests/limits for GPUs:
nvidia.com/gpu: 1
For finer sharing:
- MIG on supported NVIDIA GPUs for partitioning one GPU into slices
- Time-slicing if acceptable, though it’s less isolated
6) Build the storage layer
Training and inference need different storage patterns.
Training storage
- Large datasets in object storage or distributed storage
- Fast local cache on nodes if possible
- Checkpoints saved to durable storage
- Common choices:
- S3 / GCS / Azure Blob
- NFS for smaller setups
- Ceph / Lustre / BeeGFS for bigger clusters
Inference storage
- Model artifacts in object storage or model registry
- Fast startup by preloading models onto local disk or using a model cache
- Keep images and models versioned
A model registry is very useful to manage versioning and promotion.
7) Set up networking
For training:
- Multi-node training benefits from low-latency, high-bandwidth networking
- Ensure NCCL works correctly
- Consider RDMA/InfiniBand for scale
For inference:
- Put an ingress or API gateway in front
- Use load balancing and possibly autoscaling
- Ensure observability on latency and error rates
8) Use the right software stack
Training stack
- PyTorch or JAX or TensorFlow
- Distributed training:
- PyTorch DDP / FSDP
- DeepSpeed
- Ray Train
- Job submission:
- Kubeflow Training Operator
- Volcano
- Slurm job scripts
Inference stack
- vLLM
- NVIDIA Triton Inference Server
- TensorRT-LLM
- TorchServe
- BentoML
- FastAPI for custom serving
- For LLMs, vLLM or Triton is often a strong choice
9) Add autoscaling
Training
- Usually scale by queue depth or submitted jobs
- Cluster autoscaler can add nodes when jobs are pending
Inference
- Scale based on:
- request rate
- latency
- GPU utilization
- queue length
For inference, you may want:
- multiple replicas per model
- warm pools of GPUs
- canary deployments for new model versions
10) Put observability in place
Track:
- GPU utilization
- GPU memory usage
- temperature and power
- node health
- network throughput
- training metrics
- inference latency, throughput, error rate
- cost per job/request
Useful tools:
- Prometheus + Grafana
- NVIDIA DCGM exporter
- Loki / ELK for logs
- OpenTelemetry for tracing
11) Secure the cluster
- Use separate namespaces or projects
- RBAC and least privilege
- Private container registry
- Secrets management:
- Kubernetes Secrets, Vault, cloud KMS
- Image scanning and signed images if possible
- Network policies between training and inference services
12) Operational recommendations
- Start with one small production inference pool and one training pool
- Keep training jobs preemptible if cost matters
- Keep inference nodes reserved for production traffic
- Use spot/preemptible instances for non-critical training
- Standardize container images for reproducibility
- Automate:
- node provisioning
- driver install
- model deployment
- job submission
Example practical setup
A simple, solid setup:
- Kubernetes cluster
- Node pools:
- 2–4 GPU inference nodes with L4 or L40S
- 4–16 GPU training nodes with A100/H100 or similar
- NVIDIA GPU Operator
- vLLM or Triton for inference
- PyTorch + FSDP/DeepSpeed for training
- S3-compatible storage for datasets and models
- Prometheus/Grafana + DCGM exporter
- Cluster autoscaler
- MIG if supported and useful
- Taints/tolerations to isolate workloads
Rule of thumb
If you want the shortest path:
- Use Kubernetes
- Create separate node pools for training and inference
- Install NVIDIA GPU Operator
- Use object storage for datasets and models
- Deploy vLLM/Triton for inference
- Use PyTorch Distributed/DeepSpeed for training
- Add autoscaling and monitoring
If you want, I can also give you:
- a reference architecture diagram
- a Kubernetes-based setup checklist
- a Slurm-based alternative
- or a cost-conscious design for a small team