Prompt
Building an inference platform on GPU cloud
Latest observation
Here’s a practical blueprint for building an inference platform on GPU cloud. I’ll keep it architecture-first and implementation-oriented.
1) Define the workload
Before choosing tools, pin down:
- Model type: LLM, vision, speech, embeddings, reranker, multimodal
- Serving pattern: real-time, batch, streaming, async jobs
- Latency target: p50 / p95 / p99
- Throughput target: requests/sec or tokens/sec
- Context sizes: input/output token ranges
- Concurrency: number of simultaneous requests
- SLA: uptime, failover, scaling expectations
- Cost target: $/1k requests, $/1M tokens, etc.
This determines whether you optimize for:
- lowest latency,
- highest throughput,
- lowest cost,
- or a hybrid.
2) Reference architecture
A solid GPU inference platform usually looks like this:
Client/API Gateway → Auth / Rate limiting → Request router → Queue / scheduler → Inference workers on GPU nodes → Post-processing → Cache / storage → Observability
Core components
- API layer: REST/gRPC
- Router/Scheduler: assigns requests to healthy model replicas
- Inference runtime: Triton, vLLM, TGI, TensorRT-LLM, Ray Serve, BentoML, KServe
- GPU nodes: autoscaled pool of instances
- Storage: model weights, artifacts, logs
- Monitoring: metrics, traces, logs
- Queueing: for async workloads or load shedding
- Cache: prompt/result cache, KV cache optimization where applicable
3) Pick the serving stack
For LLMs
Common choices:
- vLLM: excellent throughput, paged attention, strong default choice
- TensorRT-LLM: best performance when models are optimized for NVIDIA GPUs
- Hugging Face TGI: easy deployment, good production experience
- NVIDIA Triton: very flexible, multi-framework serving
- Ray Serve: good orchestration and scaling, often paired with vLLM/Triton
For vision / classic ML
- Triton Inference Server
- TorchServe (less common now)
- KServe/BentoML wrappers around your model container
Rule of thumb
- Want fast LLM deployment with strong throughput? vLLM
- Want maximum performance on NVIDIA stack? TensorRT-LLM
- Want multi-model/multi-framework serving? Triton
- Want platform simplicity and Kubernetes integration? KServe + vLLM/Triton
4) Infra foundation
Cloud/GPU node setup
- Use a dedicated GPU node pool
- Keep CPU-only control plane
- Separate:
- prod
- staging
- offline/batch
- Prefer instance types based on model size:
- smaller models: L4, T4, A10
- larger throughput: A100, H100, H200
- Track:
- GPU memory
- compute utilization
- PCIe/NVLink characteristics
- network bandwidth
Kubernetes vs bare metal
Kubernetes is usually best if you need:
- autoscaling
- rolling deploys
- multi-tenant workloads
- observability and policy
Bare metal / managed VM pools can work if:
- you want simpler ops
- you have a single-purpose inference fleet
- you’re optimizing aggressively for cost/latency
A common practical choice: Kubernetes + GPU node autoscaling.
5) Model packaging and deployment
Best practices
- Package each model as a versioned container image
- Separate:
- model code
- model weights
- runtime dependencies
- Store weights in:
- object storage (S3/GCS/Azure Blob)
- model registry (MLflow, Hugging Face Hub, internal registry)
Deployment flow
- Build image
- Run validation tests
- Download weights on startup or mount via cache
- Warm up the model
- Register health checks
- Gradually shift traffic
Important
Keep startup time low:
- pre-download weights if possible
- use local NVMe cache
- avoid heavyweight initialization in the request path
6) Request routing and batching
This is critical for GPU efficiency.
Techniques
- Dynamic batching: combine multiple requests into a single GPU call
- Micro-batching: short window batching to reduce latency impact
- Token-aware scheduling: especially for LLMs
- Priority queues: separate latency-sensitive and background traffic
- Admission control: reject or defer requests when saturated
For LLMs
Use a serving engine that supports:
- continuous batching
- paged attention / KV cache management
- prefill vs decode optimization
This is often the difference between mediocre and excellent GPU utilization.
7) Autoscaling strategy
You usually need two layers:
A. Pod-level autoscaling
Scale inference replicas based on:
- GPU utilization
- queue depth
- tokens/sec per replica
- latency SLOs
- active sequences
B. Node-level autoscaling
Add/remove GPU nodes based on pending pods or actual utilization.
Practical metrics
For LLMs, don’t rely only on CPU or generic GPU percent. Track:
- active requests
- queue length
- time to first token
- tokens generated/sec
- KV cache usage
- OOM events
- p95 latency
8) Reliability and rollout
Deployment strategies
- Canary
- Blue/green
- Shadow traffic
- Gradual percentage rollout
Must-haves
- readiness/liveness probes
- request timeouts
- retry policy with backoff
- circuit breakers
- fallback model or degraded mode
- graceful shutdown to finish in-flight requests
Model versioning
Treat model versions like software releases:
- model name + version
- runtime version
- tokenizer version
- prompt template version
- feature flags
9) Caching
Caching saves huge amounts of money.
Useful caches
- Prompt/result cache for repeated prompts
- Embedding cache
- Tokenization cache
- Metadata/config cache
For LLMs
- Cache common system prompts
- Reuse prompt prefixes where possible
- Consider semantic caching for repeated user queries
Be careful with cache invalidation and privacy.
10) Observability
You need observability at 3 layers:
Infrastructure
- GPU utilization
- GPU memory usage
- node health
- network I/O
- disk I/O
Model serving
- request count
- queue time
- prefill/decode latency
- tokens/sec
- batch size distribution
- error rate
- OOMs
- timeouts
Product
- user latency
- completion quality
- abandonment rate
- cost per request
- inference success rate
Tools
- Prometheus + Grafana
- OpenTelemetry
- Loki/ELK for logs
- tracing for request-level debugging
11) Cost optimization
This is often the hardest part.
Common optimizations
- Right-size GPU instances
- Use quantization: FP16/BF16, INT8, 4-bit where acceptable
- Use tensor/pipeline parallelism only when needed
- Increase batching
- Trim context lengths
- Use smaller models for routing / classification
- Spot instances for non-critical or batch workloads
- Scale-to-zero for idle environments if acceptable
Big levers for LLMs
- reduce average input tokens
- reduce output tokens
- use speculative decoding
- use smaller draft models
- route easy requests to smaller models
12) Security and governance
- IAM least privilege
- encrypt data at rest and in transit
- private networking/VPC peering
- secrets management
- audit logs
- tenant isolation if multi-tenant
- PII redaction / policy filters
- model and prompt access controls
If you handle regulated data, define:
- data retention
- deletion policy
- locality requirements
- compliance logging
13) Recommended stack examples
Simple LLM platform
- Kubernetes
- vLLM
- NVIDIA L4/A10/A100
- Prometheus/Grafana
- S3 for model storage
- Ingress/API gateway
- Redis for caching
- HPA/Karpenter for autoscaling
High-performance enterprise stack
- Kubernetes
- TensorRT-LLM or Triton
- H100/A100
- Ray or custom scheduler
- OpenTelemetry + Prometheus
- Blue/green deployment
- Semantic cache + rate limits
- Private VPC + strict IAM
14) Build order
If you’re starting from scratch, I’d build in this order:
- Single-model container
- Health checks and metrics
- Basic API gateway
- GPU node pool
- Autoscaling
- Dynamic batching
- Versioned rollout
- Caching
- Multi-model routing
- Cost optimization
- Fallbacks and SLO automation
15) A minimal production architecture
If you want the simplest viable production setup:
- API gateway
- Kubernetes
- vLLM inference pods
- GPU autoscaling
- Prometheus/Grafana
- S3 model storage
- Redis cache
- Canary deploys
- basic auth/rate limiting
That’s enough for a strong v1.
If you want, I can also provide one of these next:
- a concrete architecture diagram
- a Kubernetes deployment example
- a GPU instance sizing guide
- a cost optimization checklist
- a reference stack for LLM inference specifically