Prompt
I'm building an app with multiple model providers — what infra should I use?
Latest observation
If you’re building an app that talks to multiple model providers (OpenAI, Anthropic, Gemini, etc.), the best infra is usually a provider-agnostic LLM layer with a few supporting pieces around it.
Recommended stack
1. LLM gateway / abstraction layer
Use a middleware layer that standardizes:
- auth
- request/response schema
- retries/failover
- streaming
- tool/function calling
- logging and tracing
- rate limits and quotas
Good options:
- LiteLLM (very common for multi-provider routing)
- OpenRouter if you want one API for many models
- Your own thin adapter if you want maximum control
Why: It prevents your app from being tightly coupled to one provider’s SDK or response shape.
2. Backend service for orchestration
Put all model calls behind your backend, not directly in the client.
Typical responsibilities:
- choose provider/model based on task, cost, latency, or availability
- manage auth keys securely
- enforce usage policies
- do prompt assembly
- run retries/timeouts/fallbacks
- post-process results
Use:
- FastAPI / Flask / Node.js / Go
- containerized with Docker
- deployed on Kubernetes, ECS/Fargate, Cloud Run, Fly.io, or Render
3. Queue + workers for async jobs
For longer tasks like:
- document processing
- batch summarization
- embeddings generation
- eval runs
- tool-heavy agent workflows
Use:
- Redis + Celery / RQ / BullMQ
- or cloud queues like SQS, Pub/Sub, RabbitMQ
Why: Keeps your API responsive and lets you retry safely.
4. Observability / tracing
This matters a lot once you’re using multiple providers.
Track:
- prompt/version
- model used
- latency
- token usage
- cost
- error rates
- output quality
- fallback frequency
Tools:
- Langfuse
- Helicone
- Arize Phoenix
- OpenTelemetry + your observability stack
5. Routing / fallback logic
Implement a policy engine that can route by:
- task type
- cost ceiling
- latency target
- context window
- provider uptime
- user tier
Example:
- cheap model for classification
- stronger model for reasoning
- fallback to another provider if primary fails
- use smaller models for autocomplete / extraction
6. Data layer
You’ll usually want:
- Postgres for app data, prompt configs, audit logs
- Redis for caching, rate limits, ephemeral state
- Object storage like S3/GCS for files, transcripts, artifacts
- optional vector DB if you’re doing RAG:
- pgvector
- Pinecone
- Weaviate
- Qdrant
7. Secrets and key management
Never store provider API keys in the frontend.
Use:
- environment variables for small deployments
- AWS Secrets Manager, GCP Secret Manager, Vault, or 1Password Secrets Automation
A good default architecture
Client
→ Your API backend
→ LLM gateway/adapter (LiteLLM or custom)
→ Model providers
With:
- Postgres for state
- Redis for cache/rate limits
- Queue/workers for async tasks
- Langfuse/Helicone for tracing
My practical recommendation
If you want fastest time-to-market:
- Backend: FastAPI or Node.js
- LLM abstraction: LiteLLM
- DB: Postgres
- Cache/queue: Redis
- Tracing: Langfuse or Helicone
- Deployment: Cloud Run / ECS / Fly.io
If you want maximum control and portability:
- build your own provider adapter layer
- keep a clean internal interface like:
generateText()streamChat()embed()moderate()
- add routing and observability yourself
What I’d avoid
- Calling providers directly from the frontend
- Hardcoding provider-specific request formats everywhere
- No tracing/cost tracking
- No fallback strategy
- Mixing orchestration logic into UI code
If you want, I can suggest a specific infra stack based on:
- your cloud (AWS/GCP/Azure/Vercel/etc.),
- expected traffic,
- whether you need RAG / agents / evals,
- and your language/framework.
Brands