Prompt
I'm building an AI workflow and want to swap between OpenAI, Anthropic, and open-source models
Latest observation
Yes — the cleanest way to do this is to add a provider abstraction layer so your workflow talks to one internal interface, while adapters handle OpenAI, Anthropic, and open-source models.
Recommended approach
1) Define one common interface
Create a single function/class your app uses everywhere, for example:
generate_text(prompt, model, provider, temperature, tools, max_tokens)- or
llm.chat(messages, options)
Then implement provider-specific adapters under the hood.
2) Normalize inputs/outputs
Different providers vary in:
- chat format
- tool/function calling
- streaming
- token limits
- system prompt handling
- response metadata
Convert everything into a shared internal schema like:
{
"messages": [
{"role": "system", "content": "..."},
{"role": "user", "content": "..."}
],
"temperature": 0.7,
"max_tokens": 1000
}
And normalize outputs to:
{
"text": "...",
"tool_calls": [],
"usage": {"input_tokens": 123, "output_tokens": 456},
"raw": {}
}
3) Use an adapter per provider
Example structure:
providers/openai.pyproviders/anthropic.pyproviders/local.py
Each adapter converts your internal schema into the provider’s API.
4) Keep model selection config-driven
Instead of hardcoding models, store them in config:
default_provider: openai
models:
fast: gpt-4.1-mini
smart: claude-3-5-sonnet-latest
cheap: llama3.1:8b
5) Add a routing layer if needed
You can route by:
- cost
- latency
- task type
- context length
- availability/fallback
Example:
- use OpenAI for structured tool use
- Anthropic for long reasoning tasks
- open-source for cheap bulk jobs or on-prem deployment
If you want the easiest implementation
Use an orchestration framework that already supports multiple providers:
- LiteLLM — very useful for switching among OpenAI, Anthropic, Azure, Bedrock, open-source gateways
- LangChain
- LlamaIndex
- Haystack
LiteLLM is especially good if your main goal is one OpenAI-like API across many models.
A practical architecture
Your app
-> LLM Router
-> OpenAI adapter
-> Anthropic adapter
-> Local model adapter (vLLM / Ollama / TGI)
Local/open-source options
For open-source models, common serving layers are:
- Ollama for local dev
- vLLM for high-throughput serving
- Text Generation Inference (TGI)
- LM Studio for easy experimentation
If you want a single HTTP API across providers, put a gateway in front.
Important compatibility issues to plan for
Tool calling
Not all providers expose it identically. You may need a unified internal representation like:
{
"name": "search_docs",
"arguments": {"query": "..." }
}
Streaming
Some providers stream deltas differently. Normalize stream events in your adapter.
JSON mode / structured output
OpenAI, Anthropic, and open-source models differ in reliability. You may want:
- schema validation
- retry on invalid JSON
- post-processing repair
Context length
Different models have different maximum context windows; your router should know this.
System prompt behavior
Some models treat system messages differently, so keep prompts tested per provider.
Best practice recommendation
If you’re early in the build:
- Define a provider-agnostic internal LLM API
- Use LiteLLM or your own adapter layer
- Add config-based model routing
- Implement fallback to another provider on failure
- Validate outputs with schemas
If you want, I can give you:
- a Python implementation of a provider-agnostic LLM wrapper,
- a LangChain/LiteLLM version, or
- a production architecture for multi-provider model routing.