Prompt
I'm building a model serving platform on Kubernetes, what should I use?
Latest observation
For a Kubernetes-based model serving platform, the best choice depends on what you need most: simple deployment, high throughput, multi-model management, or full MLOps integration. Here’s the practical breakdown.
Best default choice
KServe
Use KServe if you want:
- Kubernetes-native model serving
- Autoscaling (including scale-to-zero)
- Standardized inference APIs
- Support for multiple runtimes
- Easy integration with Istio/Knative, model storage, and canary rollout patterns
Why it’s a strong default:
KServe is one of the most complete Kubernetes-native model serving frameworks. It’s good for production platforms where you want a clean abstraction over different model backends.
If you need max GPU inference performance
NVIDIA Triton Inference Server
Use Triton if your priority is:
- High-performance GPU serving
- Dynamic batching
- Concurrent model execution
- TensorRT / PyTorch / TensorFlow / ONNX support
- Low latency and high throughput
Common pattern:
Use KServe as the platform layer and Triton as the model runtime behind it.
If you want flexible Python-first serving
BentoML
Use BentoML if you want:
- Fast setup
- Python-native model packaging
- Custom preprocessing/postprocessing
- Easy APIs and deployment workflows
- Good developer experience
This is often a good choice for teams moving quickly, especially if they have custom logic around the model.
If you want lightweight, general-purpose model hosting
Seldon Core
Use Seldon Core if you need:
- Kubernetes-native deployment
- Advanced inference graphs / pipelines
- Model explainability integrations
- Multi-step inference workflows
It’s useful, but in many teams KServe has become the more common default for new Kubernetes model serving platforms.
If you want a full internal platform with pipelines and experiment tracking
MLflow + serving layer
Use MLflow for:
- Model registry
- Lifecycle management
- Experiment tracking
Then pair it with:
- KServe for serving
- or Triton for high-performance GPU inference
MLflow is not really a serving platform by itself, but it’s very useful as the system of record for models.
My recommendation by scenario
1. General-purpose production platform
Choose: KServe + MLflow
- KServe for deployment/serving
- MLflow for registry and model lineage
- Prometheus/Grafana for monitoring
- Istio or Knative depending on your traffic patterns
2. GPU-heavy, latency-sensitive serving
Choose: Triton + KServe
- Triton for runtime performance
- KServe for Kubernetes abstraction and traffic management
3. Small team, fastest time to value
Choose: BentoML
- Simpler operational overhead
- Easier app-style model packaging
- Good for custom inference code
4. Complex inference pipelines or graph-based routing
Choose: Seldon Core or custom services
- Especially if you need ensembles, chains, or explainability hooks
A sensible platform architecture
A common production setup is:
- KServe for deployment and autoscaling
- Triton for GPU model execution
- MLflow for registry/tracking
- Prometheus + Grafana for observability
- Istio/Knative for traffic routing and scale-to-zero
- Argo CD / Flux for GitOps deployment
Short answer
If you want one recommendation:
Use KServe as your serving platform.
If you’re serving GPU models or care a lot about performance, use Triton behind KServe.
If you want, I can also give you:
- a reference architecture for a Kubernetes model serving platform,
- a feature comparison table of KServe vs Triton vs BentoML vs Seldon, or
- a suggested production stack based on your model types and traffic patterns.