Prompt
What should I use for a serverless-ish inference layer?
Latest observation
If you want a “serverless-ish” inference layer, the best choice depends mostly on model size, latency needs, and whether you want to run your own model or use a hosted API.
Good options by use case
1) Managed model APIs — easiest “serverless”
Use these if you want almost no ops:
- OpenAI / Anthropic / Google Gemini / Mistral API
- Best for: chat, embeddings, extraction, moderation, agent workflows
- Pros: autoscaling, no infra, strong reliability
- Cons: per-token cost, less control, vendor dependency
If your app can use an external model, this is usually the cleanest serverless-ish layer.
2) Serverless GPU inference platforms
Use these if you need to host your own models but still want scale-to-zero or minimal ops:
- Modal
- Runpod Serverless
- Replicate
- Baseten
- Together AI hosted endpoints
- Hugging Face Inference Endpoints (more managed than serverless, but close)
Best for: custom open-source models, fine-tunes, image/audio models, bursty traffic
3) Cloud serverless containers with autoscaling
If you want more control over the runtime:
- AWS Lambda for smaller CPU models or pre/post-processing
- Cloud Run (very popular for “serverless-ish” inference)
- AWS ECS/Fargate
- Azure Container Apps
- Google Cloud Run
Best for: lighter models, wrapper services, orchestration around a separate model host
4) Specialized inference serving stacks
If you’re willing to manage a bit more:
- vLLM
- TGI (Text Generation Inference)
- SGLang
- Ray Serve
- KServe / BentoML
- NVIDIA Triton for broader ML serving
Best for: efficient self-hosted LLM serving, batching, throughput optimization
My practical recommendation
- Fastest path: use a hosted model API
- Need your own model: use Modal or Runpod Serverless
- Want cloud-native and simple: use Cloud Run
- Need serious throughput on self-hosted LLMs: use vLLM on a managed GPU service
Rule of thumb
- Small/medium traffic + no infra team → hosted API
- Custom model + bursty traffic → Modal / Runpod / Baseten
- Need private deployment → Cloud Run / Fargate / dedicated GPU endpoint
- High-throughput LLM serving → vLLM-based endpoint
If you tell me:
- model type (LLM, embedding, vision, audio),
- expected traffic,
- latency target,
- whether you need scale-to-zero,
…I can recommend a concrete stack.