Prompt
How do I normalize different LLM APIs so my app can swap providers easily?
Latest observation
A good way to make LLM providers swappable is to define one internal interface in your app, then build a provider adapter for each API. Your app talks only to your interface; adapters translate to OpenAI, Anthropic, Gemini, etc.
1) Pick a canonical internal model
Decide what your app needs to support, and normalize everything around that:
- Input
- system / developer / user / assistant messages
- tools or function specs
- image/audio/file inputs if needed
- generation settings:
temperature,top_p,max_tokens,stop, etc.
- Output
- text
- tool calls
- usage metadata
- finish reason
- provider-specific IDs / latency / raw response for debugging
Keep this schema as your app’s “truth.”
2) Use an adapter interface
Define a provider-agnostic contract, for example:
interface LLMClient {
chat(request: ChatRequest): Promise<ChatResponse>;
streamChat?(request: ChatRequest): AsyncIterable<ChatChunk>;
}
Where ChatRequest might look like:
type Role = "system" | "user" | "assistant" | "tool";
interface Message {
role: Role;
content: string | ContentPart[];
name?: string;
}
interface ChatRequest {
model: string;
messages: Message[];
temperature?: number;
topP?: number;
maxTokens?: number;
stop?: string[];
tools?: ToolSpec[];
toolChoice?: "auto" | "none" | { name: string };
responseFormat?: { type: "json_schema" | "text"; schema?: unknown };
}
And ChatResponse:
interface ChatResponse {
text?: string;
toolCalls?: Array<{
id?: string;
name: string;
arguments: unknown;
}>;
usage?: {
inputTokens?: number;
outputTokens?: number;
totalTokens?: number;
};
finishReason?: "stop" | "length" | "tool_calls" | "content_filter" | "error";
raw?: unknown;
}
3) Translate provider quirks in adapters
Each provider has different names and capabilities. The adapter handles:
- message format differences
- some support
system - some prefer
instructions - some have separate “developer” roles
- some support
- tool calling differences
- JSON schema vs plain function definitions
- one tool call vs multiple tool calls
- token/usage differences
- fields may be named differently or missing
- streaming differences
- chunk shapes vary a lot
Your app should never know these details.
4) Normalize capabilities, not just payloads
Providers differ in what they can do. Expose a capability matrix:
interface ProviderCapabilities {
text: boolean;
streaming: boolean;
toolCalling: boolean;
jsonMode: boolean;
vision: boolean;
maxContextTokens?: number;
}
Before sending a request, either:
- degrade gracefully, or
- reject unsupported features early with a clear error
Example:
- If provider doesn’t support tool calling, you can emulate with JSON prompting or disable that route.
- If provider doesn’t support structured outputs, fall back to “return valid JSON only” prompts.
5) Keep prompt construction separate from provider adapters
A clean architecture is:
- Prompt builder: turns app state into canonical messages
- Router: chooses provider/model
- Adapter: converts canonical request to provider API
- Normalizer: converts provider response back to canonical response
This avoids tangled logic.
6) Normalize errors too
Different APIs fail differently. Convert them into a shared error model:
class LLMError extends Error {
code: "auth" | "rate_limit" | "timeout" | "invalid_request" | "provider_error";
provider?: string;
status?: number;
retryable?: boolean;
raw?: unknown;
}
This lets your app do consistent retry/backoff, fallback, and user messaging.
7) Preserve raw responses for debugging
Even with normalization, store:
- raw request
- raw response
- provider name
- model name
- request ID
- latency
This is invaluable when a provider behaves oddly.
8) Handle streaming with a unified event model
Streaming APIs differ a lot, so normalize to events:
type ChatEvent =
| { type: "start" }
| { type: "delta_text"; text: string }
| { type: "delta_tool_call"; name?: string; argumentsDelta?: string }
| { type: "end"; usage?: ChatResponse["usage"]; finishReason?: string }
| { type: "error"; error: LLMError };
Your UI only listens to this event stream.
9) Build a provider registry
Use configuration to select providers:
const providers = {
openai: new OpenAIAdapter(...),
anthropic: new AnthropicAdapter(...),
gemini: new GeminiAdapter(...),
};
Then route by:
- explicit user choice
- environment/config
- fallback order
- latency/cost policy
- feature requirement
10) Version your internal schema
As your app grows, your canonical interface will evolve. Version it so adapters don’t break silently.
Practical recommendation
If you want the simplest robust design:
- Define a canonical chat schema
- Implement one adapter per provider
- Normalize responses, usage, errors, and streaming
- Add a capability registry
- Keep raw payloads for observability
Tiny example
Your app calls:
await llm.chat({
model: "latest",
messages: [
{ role: "system", content: "You are a helpful assistant." },
{ role: "user", content: "Summarize this text..." }
],
temperature: 0.2,
tools: [{ name: "search", description: "Search docs", parameters: { type: "object", properties: { q: { type: "string" } } } }]
});
The OpenAI adapter converts to OpenAI format, Anthropic adapter converts to Anthropic format, and both return the same internal ChatResponse.
If you want, I can also give you:
- a TypeScript implementation template
- a Python interface + adapters
- or a provider comparison map for OpenAI / Anthropic / Gemini / Mistral.