Prompt
I'm tired of fighting EKS for model serving, what should I switch to?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
If your main goal is serving models reliably with less cluster pain, I’d usually recommend moving off EKS unless you specifically need Kubernetes-level control.
Best switch depends on what you’re serving
1) Simplest managed path: AWS SageMaker
Best if you want:
- managed autoscaling
- easy model deployment/versioning
- GPU inference without managing nodes
- built-in monitoring/invocation logging
Good for:
- standard online inference
- common ML stacks
- teams that don’t want to own infra
Tradeoff:
- can get expensive
- some lock-in
- less flexibility than raw containers/K8s
2) If you want “just run containers”: ECS on EC2 or Fargate
Best if your serving stack is already containerized and you want:
- less operational complexity than EKS
- easier networking/service discovery than K8s
- no control plane gymnastics
Good for:
- REST/gRPC model APIs
- moderate traffic
- simpler microservice-style serving
Tradeoff:
- not as feature-rich for ML-specific workflows as SageMaker
- GPU support is possible, but not as seamless as a dedicated ML platform
3) If you need scale-to-zero / spiky traffic: Cloud Run / serverless-style platforms
If your workload is bursty and you can tolerate cold starts:
- Cloud Run if you’re on GCP
- Azure Container Apps if you’re on Azure
- on AWS, similar patterns exist but are less clean for GPU-heavy model serving
Good for:
- lightweight models
- low/variable request volume
- teams prioritizing simplicity over max throughput
Tradeoff:
- GPUs and large models can be awkward
- cold start latency
4) If you want maximum control and fewer abstractions: plain EC2 + systemd + a load balancer
Old-school, but often the least annoying for model serving if:
- you have a small number of endpoints
- you want predictable performance
- you’re tired of orchestration overhead
Common setup:
- one AMI or container image
- systemd or Docker Compose
- ALB/NLB in front
- autoscaling group behind it
Good for:
- latency-sensitive inference
- GPU workloads
- small teams that prefer explicit infrastructure
Tradeoff:
- you own everything
- less elegant rollout/rollback story unless you build it
My practical recommendation
If you’re “tired of fighting EKS,” I’d choose based on how much you value control vs. convenience:
- Pick SageMaker if you want the least ops burden for ML inference.
- Pick ECS if you want to keep containers but ditch Kubernetes.
- Pick EC2 + LB if performance/control matters more than platform features.
- Pick a serverless container platform if traffic is spiky and models are small enough.
One rule of thumb
If your model serving stack is mostly:
- Python API + model file + autoscaling → SageMaker or ECS
- GPU-heavy or low-latency custom serving → EC2 or SageMaker
- tiny team and no appetite for platform work → SageMaker
- want to stay in AWS but avoid K8s entirely → ECS first
If you want, I can give you a direct recommendation based on your model size, traffic pattern, GPU needs, and latency target.