Prompt
Why am I stuck with SageMaker and what are people using instead?
Latest observation
People usually feel “stuck” with SageMaker for one of two reasons:
- They already built on it, and there’s real migration cost.
- AWS pushes a lot of machine-learning infrastructure through SageMaker, so it becomes the default if you’re already on AWS.
Why SageMaker feels sticky
1. It’s an all-in-one platform
SageMaker covers:
- training
- tuning
- model registry
- endpoints for serving
- pipelines
- notebooks
- feature store, grounding, etc.
That convenience is also the trap: once your workflows, IAM roles, VPC setup, artifacts, and deployment scripts are built around it, moving pieces out gets annoying.
2. AWS integration is deep
If your data lives in:
- S3
- Redshift
- Glue
- Athena
- ECR
- IAM/VPC/KMS
then SageMaker is the path of least resistance. The setup feels “natural” in AWS, but it also means the rest of your ML stack gets coupled to AWS-specific primitives.
3. Managed, but opinionated
SageMaker reduces ops work, but not always in the way teams want. Common pain points:
- confusing pricing
- endpoint management overhead
- notebook lifecycle issues
- SDK complexity
- lots of AWS-specific abstractions
- less portability than plain Docker/Kubernetes/standard Python tooling
4. Portability gets worse over time
A lot of teams start with “we’ll just use SageMaker for convenience,” then later realize:
- training scripts depend on SageMaker job APIs
- deployment assumes SageMaker endpoints
- pipelines are encoded in AWS-specific services
- notebooks and artifacts are tied to AWS accounts and roles
That makes switching expensive even if the technical alternative is simpler.
What people are using instead
1. Plain Docker + Kubernetes
Very common for teams that want control and portability.
Typical stack:
- Docker for packaging
- Kubernetes for orchestration
- Argo Workflows / Kubeflow / Flyte / Dagster / Prefect for pipelines
- KServe / Seldon / BentoML / Ray Serve for serving
Why people choose it:
- portable across clouds
- more transparent
- less vendor lock-in
- easier to standardize with general platform engineering
Tradeoff:
- more ops burden than a managed platform
2. Databricks
A lot of organizations use Databricks instead of SageMaker, especially if the core problem is:
- feature engineering
- distributed training
- data/ML workflows
- notebooks for analysts and ML engineers
Why people like it:
- strong Spark/data ecosystem
- easier collaboration
- unified data + ML workflows
- decent governance and lifecycle tooling
Tradeoff:
- not always ideal for low-latency model serving
- still a platform lock-in, just a different one
3. Open-source ML platforms
Common choices:
- MLflow for experiment tracking/model registry
- Kubeflow for ML pipelines
- Airflow / Dagster / Prefect for orchestration
- BentoML for model packaging/serving
- Ray for distributed training/inference
This is popular with teams that want modular tooling rather than one giant platform.
Tradeoff:
- you assemble and maintain the stack yourself
4. Vertex AI or Azure ML
If the reason for SageMaker pain is mainly “we want a managed cloud ML platform, but not AWS,” then alternatives are:
- Google Vertex AI
- Azure Machine Learning
These are conceptually similar to SageMaker:
- managed training
- deployment
- registry
- pipelines
People often switch for:
- better integration with their cloud/vendor choice
- simpler UX in some areas
- internal company standardization
Tradeoff:
- still a cloud lock-in, just not AWS
5. Specialized serving tools
If your main use of SageMaker is inference rather than training, people often replace it with:
- BentoML
- KServe
- NVIDIA Triton
- Ray Serve
- TorchServe (less trendy now, but still used)
Why:
- simpler deployment model
- better performance/control
- easier to run on Kubernetes or plain containers
6. Just standard app infrastructure
For smaller models or non-heavy ML systems, teams often skip a “platform” entirely:
- train in Python scripts or notebooks
- track with MLflow
- serve via FastAPI/Flask
- containerize with Docker
- deploy on ECS, EKS, Cloud Run, or plain VMs
This is especially common when:
- model traffic is low
- latency requirements are moderate
- the team wants fewer moving parts
So why are people moving away from SageMaker?
Usually because they want one or more of these:
- less AWS lock-in
- simpler architecture
- more portable deployment
- lower cost
- better developer experience
- more control over serving and orchestration
- closer alignment with platform engineering standards
SageMaker is powerful, but some teams eventually decide it’s “too much platform” for what they need.
A practical rule of thumb
- If you’re all-in on AWS and want a managed ML platform, SageMaker can make sense.
- If you want portability and control, people often go with Docker + Kubernetes + MLflow + a workflow tool.
- If your workload is mostly data/analytics + ML, Databricks is common.
- If you mainly need simple model serving, BentoML/Ray Serve/FastAPI are popular.
If you want, I can also give you:
- a “SageMaker vs Databricks vs Kubeflow vs Vertex AI” comparison, or
- a modern reference stack people use instead of SageMaker in 2026.