Prompt

How do I normalize different LLM APIs so my app doesn't need provider-specific code?

Artificial Intelligence / AI Infrastructure1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

A good way to normalize LLM APIs is to build a provider-agnostic adapter layer inside your app.

Core idea

Define one internal interface for what your app needs, then write small adapters for each provider that translate:

  • your internal request shape → provider request shape
  • provider response shape → your internal response shape

That keeps provider-specific code isolated in one place.


1) Define a canonical internal schema

Pick the smallest set of features your app actually uses.

Example internal request:

type LLMRequest = {
  model: string
  messages: Array<{ role: 'system' | 'user' | 'assistant'; content: string }>
  temperature?: number
  maxTokens?: number
  stream?: boolean
  tools?: ToolSpec[]
}

Example internal response:

type LLMResponse = {
  text: string
  finishReason?: 'stop' | 'length' | 'tool_calls' | 'error'
  usage?: {
    inputTokens?: number
    outputTokens?: number
    totalTokens?: number
  }
  toolCalls?: Array<{
    name: string
    arguments: unknown
  }>
  raw?: unknown
}

Only include fields you truly need. Over-normalizing makes the abstraction fragile.


2) Build a provider interface

Your app talks to this interface only:

interface LLMClient {
  complete(request: LLMRequest): Promise<LLMResponse>
  stream?(request: LLMRequest): AsyncIterable<LLMResponseChunk>
}

Then implement one adapter per provider:

  • OpenAIClient
  • AnthropicClient
  • GeminiClient
  • MistralClient
  • etc.

3) Put translation logic in adapters

Each adapter handles differences such as:

  • message formats
  • system prompt support
  • tool/function calling format
  • streaming events
  • token usage fields
  • stop reasons
  • image/audio/document inputs

Example:

class OpenAIClient implements LLMClient {
  constructor(private apiKey: string) {}

  async complete(req: LLMRequest): Promise<LLMResponse> {
    const providerReq = {
      model: req.model,
      messages: req.messages,
      temperature: req.temperature,
      max_tokens: req.maxTokens,
    }

    const res = await fetch('https://api.openai.com/v1/chat/completions', {
      method: 'POST',
      headers: {
        Authorization: `Bearer ${this.apiKey}`,
        'Content-Type': 'application/json',
      },
      body: JSON.stringify(providerReq),
    })

    const json = await res.json()

    return {
      text: json.choices?.[0]?.message?.content ?? '',
      finishReason: json.choices?.[0]?.finish_reason,
      usage: {
        inputTokens: json.usage?.prompt_tokens,
        outputTokens: json.usage?.completion_tokens,
        totalTokens: json.usage?.total_tokens,
      },
      raw: json,
    }
  }
}

4) Normalize the tricky parts explicitly

These areas usually break “universal” abstractions:

Messages

Providers differ on:

  • system message handling
  • multipart content
  • assistant tool call messages

Solution: define your own canonical message format, then map per provider.

Tool/function calling

Different providers represent tools differently. Normalize to something like:

type ToolSpec = {
  name: string
  description?: string
  inputSchema: object
}

And normalize tool calls in responses:

type ToolCall = {
  id?: string
  name: string
  arguments: unknown
}

Streaming

Some providers stream tokens, others stream deltas, others stream events.

Solution: normalize to one event stream:

type LLMResponseChunk =
  | { type: 'text'; delta: string }
  | { type: 'tool_call'; toolCall: ToolCall }
  | { type: 'done'; usage?: ... }

5) Handle capability differences

Not every provider supports every feature. Add a capability descriptor:

type LLMCapabilities = {
  chat: boolean
  streaming: boolean
  tools: boolean
  vision: boolean
  jsonMode: boolean
}

Each adapter exposes what it supports, so your app can degrade gracefully.


6) Separate configuration from behavior

Don’t scatter provider settings throughout your code. Instead keep:

  • provider name
  • model name
  • API key
  • base URL
  • timeouts
  • retries

in a single config layer.

Example:

type ProviderConfig = {
  provider: 'openai' | 'anthropic' | 'gemini'
  apiKey: string
  model: string
}

Then a factory chooses the right adapter.


7) Use an internal “LLM gateway”

A common architecture:

App → LLM Gateway/Facade → Provider Adapters → Providers

This lets you centralize:

  • retries
  • rate limiting
  • logging
  • prompt templates
  • fallbacks
  • cost tracking
  • caching
  • eval hooks

8) Don’t normalize everything

Good abstractions keep your app portable without hiding important differences.

Avoid forcing all providers into a fake identical model if:

  • some support multimodal input and others don’t
  • some have reasoning tokens, some don’t
  • some require special safety settings
  • some have different tool semantics

Instead:

  • normalize the common case
  • expose provider-specific escape hatches when needed

Example:

type LLMRequest = {
  ...
  providerOptions?: Record<string, unknown>
}

9) Consider using an existing abstraction library

If you don’t want to build everything yourself, look at libraries/frameworks that already normalize provider APIs, such as:

  • Vercel AI SDK
  • LangChain
  • LiteLLM
  • Haystack / other orchestration layers

These can reduce boilerplate, though you may still want your own thin internal interface.


Recommended practical approach

For most apps:

  1. Define one internal request/response schema
  2. Implement adapter classes per provider
  3. Keep streaming and tools normalized
  4. Add capability flags
  5. Put retries/logging/fallbacks in a gateway layer
  6. Allow provider-specific options as an escape hatch

Rule of thumb

If your app has to ask “which provider is this?” outside the adapter layer, your abstraction is leaking.

If you want, I can also give you:

  • a TypeScript interface + adapter skeleton
  • a Python version
  • or a comparison of OpenAI vs Anthropic vs Gemini normalization pitfalls

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.