Prompt
I'm building with Kubernetes and need a model serving setup that won't be a mess
Latest observation
If you want Kubernetes model serving that doesn’t turn into a mess, the main goal is to separate concerns and standardize the deployment pattern early.
A sane default architecture
Use this split:
-
Model artifact storage
- S3 / GCS / MinIO
- Store versioned model files there, not in images
-
Serving runtime
- One container per model server
- Common options:
- KServe if you want Kubernetes-native inference management
- Seldon if you want more mature ML workflow integrations
- Ray Serve if you need flexible Python-native composition
- BentoML if you want straightforward packaging and deployment
- Triton Inference Server if you mainly serve deep learning models and want high performance
-
Deployment layer
- Use Kubernetes Deployments for simple cases
- Use KServe/Seldon if you need:
- autoscaling
- canary deployments
- model revision management
- inference-specific CRDs
-
Traffic management
- Ingress or service mesh
- Keep routing outside the model container
- Use stable endpoints per model/version
-
Observability
- Metrics: latency, error rate, throughput, queue depth
- Logs: structured request IDs
- Tracing if you have chained inference
- Prometheus + Grafana is the usual baseline
What usually makes it a mess
Avoid these common traps:
-
Putting model files into Docker images
- Every model update becomes a new image build
- Images get huge and slow to deploy
-
Mixing training and serving
- Different resource profiles, different release cadence
- Keep them separate
-
Ad hoc Helm charts per model
- Leads to copy-paste sprawl
- Instead define a reusable template
-
No versioning strategy
- You want:
- model version
- runtime version
- config version
- Track them independently
- You want:
-
No clear CPU/GPU policy
- Decide up front which models can run on CPU only
- Standardize GPU node pools if needed
-
Too many bespoke sidecars
- Each extra component adds operational overhead
- Only add what you truly need
A practical pattern that works well
For most teams, this is a good default:
- Kubernetes
- KServe
- S3-compatible model store
- Prometheus/Grafana
- Istio or NGINX ingress
- GitOps with Argo CD or Flux
- CI pipeline builds runtime image only
- Model artifact promoted separately
This gives you:
- repeatable deployments
- clean versioning
- easier rollback
- less YAML drift
Simple operating model
Use this lifecycle:
- Train model
- Validate model
- Upload artifact to object storage
- Update a small config/manifest with model URI and version
- GitOps deploys serving update
- Monitor metrics
- Roll back by flipping model version
This is much cleaner than rebuilding the whole app for every model change.
If you want the least painful choice
Choose KServe if:
- you want Kubernetes-native serving
- multiple models and versions
- autoscaling / rollout control
- standardized inference APIs
Choose BentoML if:
- you want simplicity
- small team
- mostly Python inference
- you’re okay with less platform complexity
Choose Triton if:
- you serve GPU-heavy deep learning models
- performance matters a lot
- your models fit Triton’s supported formats well
Choose Ray Serve if:
- inference logic is complex
- you need multi-stage Python pipelines
- you want flexible request composition
My recommendation
If your priority is “won’t be a mess”, I’d start with:
- BentoML for simplest path, or
- KServe if you expect platform growth
And make these non-negotiable:
- model artifacts in object storage
- no model blobs in images
- GitOps for deployments
- standardized metrics/logging
- clear versioning
If you want, I can give you a reference Kubernetes architecture, or a minimal example using KServe/BentoML, or help you choose between KServe vs BentoML vs Triton for your specific models.