Prompt
We outgrew Vercel AI tooling, what do teams use for production serving?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
For production serving after outgrowing Vercel AI tooling, teams usually move to one of a few patterns depending on scale, latency, and how much they want to own.
Common production options
1) Dedicated inference providers
Best when you want to avoid running infra yourself.
- OpenAI / Anthropic / Google Gemini / Mistral APIs for hosted model access
- Together.ai, Fireworks.ai, Replicate, Groq, Cerebras for lower-level or faster model serving
- Bedrock / Vertex AI / Azure OpenAI if you want enterprise cloud-native procurement/compliance
Why teams use them:
- minimal ops
- autoscaling handled for you
- good for fast iteration
Tradeoffs:
- less control
- cost can get high at scale
- vendor-specific rate limits / latency behavior
2) Self-hosted inference on Kubernetes / GPU fleets
Best for control, cost optimization at scale, or model customization.
Typical stack:
- vLLM for high-throughput LLM serving
- TGI (Text Generation Inference) from Hugging Face
- TensorRT-LLM for optimized NVIDIA deployments
- Ray Serve or custom FastAPI/gRPC wrappers
- Kubernetes + GPU nodes + autoscaling
Why teams use them:
- can run open-weight models
- higher throughput/efficiency
- better cost control at steady high volume
Tradeoffs:
- ops complexity
- model/version management burden
- need observability, capacity planning, and fallbacks
3) Managed model serving platforms
A middle ground between APIs and self-hosting.
Examples:
- Modal
- RunPod
- Baseten
- SageMaker Endpoints
- Vertex AI Endpoints
- NVIDIA NIM (depending on deployment style)
Why teams use them:
- easier than managing GPU infra
- more control than pure APIs
- good for custom endpoints, batching, autoscaling
Tradeoffs:
- platform costs and abstractions
- portability can be limited
What production stacks often look like
For LLM apps:
- Frontend: Next.js / React
- API layer: FastAPI / Node / Go
- LLM gateway: LiteLLM, custom router, or provider abstraction
- Model serving: OpenAI/Anthropic or vLLM/TGI
- Caching: Redis
- Queueing: SQS / PubSub / Kafka / RabbitMQ
- Observability: OpenTelemetry, Prometheus, Grafana, Datadog, Langfuse, Helicone
- RAG: pgvector / Pinecone / Weaviate / Milvus / Elasticsearch
What to choose by scenario
If you want fastest path to production:
- stay on hosted APIs but add:
- your own backend
- retries/timeouts
- prompt/version management
- logging/metrics
- fallback routing
If cost is becoming the issue:
- consider vLLM on Kubernetes or managed GPU serving like Baseten/RunPod/Modal
If you need enterprise/compliance:
- Azure OpenAI, Bedrock, Vertex AI, or self-hosted in your cloud VPC
If you need high throughput/low latency:
- Groq, Fireworks, Together, or self-hosted optimized serving with batching and quantization
A very common “grown-up” architecture
Teams often end up with:
- Provider abstraction layer: LiteLLM or custom
- Primary model provider: OpenAI/Anthropic
- Fallback provider: another API or open-weight model
- Queue-based background jobs for slow tasks
- Tracing and evals: Langfuse / LangSmith / Helicone
- Canary releases for prompts/models
- Cost controls: per-user budgets, token limits, rate limiting
Practical recommendation
If you’re beyond Vercel AI SDK and want a production setup:
- Keep the app framework you like
- Move model calls behind a server-side gateway
- Add:
- timeouts/retries
- streaming support
- tracing
- caching
- fallback providers
- Decide whether to:
- stay managed: OpenAI/Anthropic + gateway
- go infra-owned: vLLM/TGI on Kubernetes
- use a managed serving platform: Modal/Baseten/RunPod
If you want, I can give you:
- a decision matrix by team size/traffic/budget, or
- a reference production architecture for serving LLMs outside Vercel.
Brands