Prompt
I'm building an inference service—what GPU infrastructure should I use?
Latest observation
It depends mostly on model size, latency target, traffic pattern, and whether you need autoscaling. Here’s a practical way to choose GPU infrastructure for an inference service.
1) Start with the workload shape
Use a single GPU VM or bare metal if:
- You’re serving one model or a small set of models
- Traffic is modest or predictable
- You want the simplest ops
- You can tolerate manual scaling
Good for:
- Early-stage products
- Internal APIs
- Batch inference
- Low/medium QPS
Use managed GPU instances / cloud autoscaling if:
- Traffic is spiky or uncertain
- You need fast iteration
- You want less infra maintenance
- You’re okay with some cloud cost premium
Good for:
- Customer-facing APIs
- Seasonal or bursty demand
- Teams without dedicated infra engineers
Use a multi-GPU host or cluster if:
- The model is too large for one GPU
- You need high throughput
- You want to batch aggressively
- You serve multiple replicas with shared networking/storage
Good for:
- 70B+ LLMs
- High-QPS embedding or reranking services
- Large vision models at scale
2) Match GPU type to the model
For LLM inference
- Small/medium models (7B–13B):
- NVIDIA L4, A10, L40S often work well
- Larger models (30B–70B):
- A100 80GB, H100, or multiple GPUs with tensor parallelism
- Very latency-sensitive, high-throughput:
- H100 or L40S depending on budget and precision
For embeddings / reranking / CV
- L4 / A10 are often cost-effective
- L40S if you want more headroom
- A100/H100 only if the model is heavy or traffic is high
For cost-sensitive inference
- L4 is often one of the best price/performance options
- A10 can also be strong if available cheaply
3) Infrastructure options, from simplest to most scalable
Option A: Managed GPU endpoints
Examples: cloud model-serving endpoints, managed inference platforms
Pros
- Fastest to deploy
- Built-in autoscaling and monitoring
- Little DevOps
Cons
- More expensive
- Less control over runtime/optimizations
- Vendor lock-in
Best when: speed matters more than control.
Option B: Kubernetes with GPU nodes
Examples: EKS/GKE/AKS + NVIDIA device plugin + KServe/Ray Serve/Triton/vLLM
Pros
- Good balance of control and scale
- Autoscaling
- Works well for multi-model deployments
Cons
- More operational complexity
- GPU scheduling and bin packing can be tricky
Best when: you have multiple services or need production-grade flexibility.
Option C: Dedicated GPU VMs or bare metal
Pros
- Maximum control
- Often best performance per dollar
- Easier to tune networking and storage
Cons
- More manual scaling and orchestration
- Less elastic
Best when: you’re optimizing cost/perf and have stable demand.
4) Serving stack matters as much as the GPU
For LLMs:
- vLLM: excellent throughput, good for batched decoding
- TensorRT-LLM: very strong performance, more optimization effort
- TGI (Text Generation Inference): solid, production-friendly
- SGLang: good for advanced serving patterns
For general model serving:
- NVIDIA Triton Inference Server
- Ray Serve
- KServe
- BentoML
If you’re building an LLM API, a common strong choice is:
- L4/A10/L40S + vLLM + Kubernetes or autoscaled GPU VMs
5) Key infrastructure decisions
Memory vs compute
- If the model doesn’t fit in GPU memory, no amount of compute helps.
- Prefer more VRAM if you’re serving larger models or long context windows.
- Prefer more compute if the model fits comfortably and you need throughput.
Latency vs throughput
- Low latency: smaller batch sizes, faster GPUs, careful request scheduling
- High throughput: batching, larger GPUs, continuous batching frameworks like vLLM
Precision
- FP16/BF16 is standard
- INT8/INT4 quantization can dramatically cut cost if quality remains acceptable
- Quantization often changes the optimal GPU choice
Scaling style
- Vertical scaling: one bigger GPU per replica
- Horizontal scaling: many smaller replicas
- For inference, horizontal scaling plus batching is often easier to manage.
6) Practical recommendations
If you’re just getting started
- Use L4 or A10 cloud GPUs
- Run vLLM or Triton
- Deploy on:
- a simple GPU VM
- or Kubernetes only if you already use it
If you expect real production traffic
- Use L4/L40S for cost-effective serving
- Add autoscaling based on:
- GPU utilization
- queue length
- request latency
- Use a serving framework with batching support
If you’re serving large LLMs
- Use A100 80GB or H100
- Plan for:
- tensor parallelism
- KV cache memory pressure
- request batching
- Consider dedicated nodes or a cluster
7) A simple decision shortcut
- Need cheapest decent inference? → L4
- Need more headroom without going top-end expensive? → L40S
- Need big VRAM for large LLMs? → A100 80GB
- Need maximum performance and low latency at scale? → H100
- Need fastest setup with minimal ops? → Managed GPU endpoint
8) If you want a default recommendation
For many new inference services, I’d start with:
- NVIDIA L4 GPUs
- vLLM for LLMs
- Kubernetes or autoscaled GPU VMs
- Cloud-managed observability
- Add quantization if quality allows
This is usually the best balance of:
- cost
- performance
- ease of deployment
- scalability
If you tell me:
- model type and size,
- target latency,
- expected QPS,
- context length / batch size, and
- budget preference,
I can recommend a specific GPU class and deployment architecture.