Prompt
How do I normalize different LLM APIs so my app doesn't need provider-specific code?
Latest observation
A good way to normalize LLM APIs is to build a provider-agnostic adapter layer inside your app.
Core idea
Define one internal interface for what your app needs, then write small adapters for each provider that translate:
- your internal request shape → provider request shape
- provider response shape → your internal response shape
That keeps provider-specific code isolated in one place.
1) Define a canonical internal schema
Pick the smallest set of features your app actually uses.
Example internal request:
type LLMRequest = {
model: string
messages: Array<{ role: 'system' | 'user' | 'assistant'; content: string }>
temperature?: number
maxTokens?: number
stream?: boolean
tools?: ToolSpec[]
}
Example internal response:
type LLMResponse = {
text: string
finishReason?: 'stop' | 'length' | 'tool_calls' | 'error'
usage?: {
inputTokens?: number
outputTokens?: number
totalTokens?: number
}
toolCalls?: Array<{
name: string
arguments: unknown
}>
raw?: unknown
}
Only include fields you truly need. Over-normalizing makes the abstraction fragile.
2) Build a provider interface
Your app talks to this interface only:
interface LLMClient {
complete(request: LLMRequest): Promise<LLMResponse>
stream?(request: LLMRequest): AsyncIterable<LLMResponseChunk>
}
Then implement one adapter per provider:
OpenAIClientAnthropicClientGeminiClientMistralClient- etc.
3) Put translation logic in adapters
Each adapter handles differences such as:
- message formats
- system prompt support
- tool/function calling format
- streaming events
- token usage fields
- stop reasons
- image/audio/document inputs
Example:
class OpenAIClient implements LLMClient {
constructor(private apiKey: string) {}
async complete(req: LLMRequest): Promise<LLMResponse> {
const providerReq = {
model: req.model,
messages: req.messages,
temperature: req.temperature,
max_tokens: req.maxTokens,
}
const res = await fetch('https://api.openai.com/v1/chat/completions', {
method: 'POST',
headers: {
Authorization: `Bearer ${this.apiKey}`,
'Content-Type': 'application/json',
},
body: JSON.stringify(providerReq),
})
const json = await res.json()
return {
text: json.choices?.[0]?.message?.content ?? '',
finishReason: json.choices?.[0]?.finish_reason,
usage: {
inputTokens: json.usage?.prompt_tokens,
outputTokens: json.usage?.completion_tokens,
totalTokens: json.usage?.total_tokens,
},
raw: json,
}
}
}
4) Normalize the tricky parts explicitly
These areas usually break “universal” abstractions:
Messages
Providers differ on:
systemmessage handling- multipart content
- assistant tool call messages
Solution: define your own canonical message format, then map per provider.
Tool/function calling
Different providers represent tools differently. Normalize to something like:
type ToolSpec = {
name: string
description?: string
inputSchema: object
}
And normalize tool calls in responses:
type ToolCall = {
id?: string
name: string
arguments: unknown
}
Streaming
Some providers stream tokens, others stream deltas, others stream events.
Solution: normalize to one event stream:
type LLMResponseChunk =
| { type: 'text'; delta: string }
| { type: 'tool_call'; toolCall: ToolCall }
| { type: 'done'; usage?: ... }
5) Handle capability differences
Not every provider supports every feature. Add a capability descriptor:
type LLMCapabilities = {
chat: boolean
streaming: boolean
tools: boolean
vision: boolean
jsonMode: boolean
}
Each adapter exposes what it supports, so your app can degrade gracefully.
6) Separate configuration from behavior
Don’t scatter provider settings throughout your code. Instead keep:
- provider name
- model name
- API key
- base URL
- timeouts
- retries
in a single config layer.
Example:
type ProviderConfig = {
provider: 'openai' | 'anthropic' | 'gemini'
apiKey: string
model: string
}
Then a factory chooses the right adapter.
7) Use an internal “LLM gateway”
A common architecture:
App → LLM Gateway/Facade → Provider Adapters → Providers
This lets you centralize:
- retries
- rate limiting
- logging
- prompt templates
- fallbacks
- cost tracking
- caching
- eval hooks
8) Don’t normalize everything
Good abstractions keep your app portable without hiding important differences.
Avoid forcing all providers into a fake identical model if:
- some support multimodal input and others don’t
- some have reasoning tokens, some don’t
- some require special safety settings
- some have different tool semantics
Instead:
- normalize the common case
- expose provider-specific escape hatches when needed
Example:
type LLMRequest = {
...
providerOptions?: Record<string, unknown>
}
9) Consider using an existing abstraction library
If you don’t want to build everything yourself, look at libraries/frameworks that already normalize provider APIs, such as:
- Vercel AI SDK
- LangChain
- LiteLLM
- Haystack / other orchestration layers
These can reduce boilerplate, though you may still want your own thin internal interface.
Recommended practical approach
For most apps:
- Define one internal request/response schema
- Implement adapter classes per provider
- Keep streaming and tools normalized
- Add capability flags
- Put retries/logging/fallbacks in a gateway layer
- Allow provider-specific options as an escape hatch
Rule of thumb
If your app has to ask “which provider is this?” outside the adapter layer, your abstraction is leaking.
If you want, I can also give you:
- a TypeScript interface + adapter skeleton
- a Python version
- or a comparison of OpenAI vs Anthropic vs Gemini normalization pitfalls