Prompt
model serving on Kubernetes
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
Model serving on Kubernetes usually means deploying an ML model as a scalable API or batch service using Kubernetes primitives like Deployments, Services, Ingress, HPA, and sometimes specialized serving platforms.
Common architecture
- Model artifact: stored in S3/GCS/Blob storage or baked into the image
- Inference container: runs a model server such as:
- TorchServe
- TensorFlow Serving
- Triton Inference Server
- FastAPI / Flask custom service
- Kubernetes Deployment: manages replicas
- Service: exposes pods internally
- Ingress / Gateway: exposes externally
- Autoscaling: HPA/KEDA based on CPU, memory, or request rate
Typical deployment flow
- Package the model server into a container image
- Push image to a registry
- Deploy to Kubernetes with a Deployment + Service
- Add readiness/liveness probes
- Add autoscaling
- Expose with Ingress or LoadBalancer
- Monitor logs, metrics, latency, and error rates
Example Deployment
apiVersion: apps/v1
kind: Deployment
metadata:
name: model-server
spec:
replicas: 2
selector:
matchLabels:
app: model-server
template:
metadata:
labels:
app: model-server
spec:
containers:
- name: server
image: your-registry/model-server:latest
ports:
- containerPort: 8080
readinessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 5
periodSeconds: 10
livenessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 15
periodSeconds: 20
Service
apiVersion: v1
kind: Service
metadata:
name: model-server
spec:
selector:
app: model-server
ports:
- port: 80
targetPort: 8080
type: ClusterIP
Autoscaling example
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: model-server-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: model-server
minReplicas: 2
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
Best practices
- Use readiness probes so traffic only goes to ready pods
- Use resource requests/limits to avoid noisy neighbors
- Keep model loading on startup optimized
- Store models outside the image for easier updates
- Use canary deployments for safer rollout
- Monitor:
- p95/p99 latency
- error rate
- GPU/CPU utilization
- model drift if applicable
- For GPU inference, use:
- node selectors / tolerations
- NVIDIA device plugin
- dedicated GPU node pools
When to use specialized tools
Use a model serving platform if you need:
- multi-model management
- advanced autoscaling
- traffic splitting / canary
- versioning
- GPU optimization
Examples:
- KServe
- Seldon Core
- BentoML
- Ray Serve
- Triton
If you want, I can give you:
- a full Kubernetes deployment example for a specific framework,
- a production-ready setup with autoscaling and ingress, or
- a comparison of KServe vs Seldon vs custom FastAPI.