Prompt
I need to serve both open-source and hosted models through one interface
Latest observation
You can do this by putting a single API gateway / abstraction layer in front of both model types, then routing requests to either:
- open-source/self-hosted models (e.g., Llama, Mistral, Qwen)
- hosted/proprietary models (e.g., OpenAI, Anthropic, Gemini, Bedrock)
Common pattern
1) Use one unified API schema
Expose one internal interface like:
POST /chatPOST /embeddingsPOST /rerankPOST /images
Then translate that request to whichever backend you want.
2) Add a model router
Route by one or more of:
- explicit model name from the client
- cost
- latency
- context length
- availability/fallback
- task type
Example:
model: "llama-3.1-70b"→ self-hosted vLLM endpointmodel: "gpt-4.1"→ OpenAImodel: "claude-sonnet-4"→ Anthropic
3) Normalize responses
Convert all vendor responses into your own standard format, e.g.:
textusagefinish_reasontool_callscitationsif applicable
4) Handle differences in capabilities
Some models support:
- function calling
- multimodal inputs
- streaming
- JSON mode
- tool use
Your interface should either:
- expose the lowest common denominator, or
- support feature flags per model and reject unsupported requests cleanly
Fastest way to implement
Option A: Use an existing proxy/router
Good if you want speed and vendor neutrality.
Examples:
- LiteLLM — very common for routing many LLM providers behind one OpenAI-compatible API
- OpenRouter — if you want a hosted aggregation layer
- Portkey — gateway, observability, routing
- LangChain / LlamaIndex — orchestration, not ideal as the main gateway
LiteLLM is especially useful if you want:
- one OpenAI-style API
- many providers
- retries/fallbacks
- cost tracking
- model aliases
Option B: Build your own thin gateway
If you want more control, make a small service that:
- accepts OpenAI-style requests
- maps model IDs to providers
- converts auth headers/secrets
- forwards the call
- normalizes the response
- logs usage and errors
This is a good fit if:
- you need custom policy
- you want strict data routing
- you need per-tenant model allowlists
- you have compliance constraints
Suggested architecture
Client
↓
Your Unified API
↓
Model Router / Policy Engine
├── Self-hosted inference (vLLM / TGI / Ollama / Triton)
├── OpenAI
├── Anthropic
└── Gemini / Bedrock / etc.
Practical recommendation
If you want the least engineering effort:
- Use LiteLLM as the gateway
- run open-source models behind vLLM or TGI
- point hosted models through their native APIs
- use a single OpenAI-compatible endpoint for your apps
If you want maximum control:
- build your own gateway and keep the API contract stable
Key design decisions to make
-
Do you want OpenAI-compatible API semantics?
- If yes, integration is much easier.
-
Do you want routing at request time or per model alias?
- Request time = flexible
- Alias-based = simpler for clients
-
Do you need fallback?
- Example: try open-source first, then hosted if overloaded
-
Do you need tenant isolation / data residency?
- This often affects provider choice and routing logic
Example model mapping
models:
gpt-4.1:
provider: openai
model: gpt-4.1
claude-sonnet:
provider: anthropic
model: claude-sonnet-4
llama-3.1-70b:
provider: vllm
endpoint: http://llm-internal:8000/v1
If you want, I can give you:
- a reference architecture
- a LiteLLM setup
- a custom FastAPI gateway example
- or a model routing policy design for your use case.