Prompt
I'm building an AI gateway and want to normalize request formats across vendors; best approach?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
A good approach is to define one internal canonical request/response schema in your gateway, then build vendor-specific adapters at the edges.
Recommended architecture
-
Canonical internal model
- Create a vendor-neutral schema for:
- model identifier
- messages / prompt
- tools / function calling
- sampling params (
temperature,top_p,max_tokens, etc.) - streaming flags
- metadata / tracing / tenant info
- Keep it stable and versioned.
- Create a vendor-neutral schema for:
-
Adapter per provider
- Convert canonical requests into each vendor’s API shape.
- Convert vendor responses back into canonical output.
- Handle provider-specific quirks here, not in business logic.
-
Capability-based normalization
- Not every vendor supports the same features.
- Maintain a capabilities registry:
- supports tools?
- supports JSON mode?
- supports logprobs?
- supports multimodal?
- supports streaming?
- Validate requests against provider capability before dispatch.
- If unsupported, either:
- reject early
- degrade gracefully
- emulate if possible
-
Typed schema + validation
- Use explicit types and strict validation for the canonical format.
- Prefer JSON Schema / OpenAPI / Pydantic / Zod / Protobuf depending on stack.
- Enforce required fields, defaults, and value ranges.
-
Preserve original payload when needed
- Support an escape hatch for vendor-specific params:
provider_optionsextensions
- This avoids blocking access to provider-only features while keeping the core normalized.
- Support an escape hatch for vendor-specific params:
What to normalize
Focus on these common areas:
- Messages
- role: system/user/assistant/tool
- content as text or structured blocks
- Tool calling
- function name, JSON args, tool call IDs
- Generation params
- temperature, top_p, max tokens, stop sequences
- Streaming
- event shape and partial deltas
- Usage
- prompt/completion tokens, total tokens
- Errors
- standardized error code, retryability, upstream provider error details
What not to over-normalize
Avoid forcing everything into one lowest-common-denominator shape if it removes useful features. Common pitfalls:
- flattening multimodal content too early
- discarding provider-specific tool metadata
- converting all errors into one generic error
- hiding model capability differences
A better pattern is:
- normalize the core
- pass through optional provider extensions
- attach provider metadata separately
Practical design pattern
A clean internal shape might look like:
{
"model": "gpt-4.1",
"messages": [
{"role": "system", "content": "You are helpful."},
{"role": "user", "content": "Hello"}
],
"tools": [],
"params": {
"temperature": 0.7,
"max_tokens": 256,
"top_p": 1.0
},
"stream": false,
"metadata": {
"tenant_id": "abc",
"request_id": "req_123"
},
"provider_options": {
"vendor_x": {
"foo": "bar"
}
}
}
Operational tips
- Version your canonical schema
- Log both canonical and translated provider payloads
- Add contract tests per provider
- Track feature parity and translation gaps
- Make retries, timeouts, and idempotency part of the gateway layer
- Keep provider adapters stateless where possible
If you’re choosing between strategies
- Single universal schema only: simple, but you’ll lose vendor-specific power.
- Pass-through with thin wrappers: flexible, but hard to maintain.
- Canonical core + provider extensions: usually best balance.
Best practice summary
The best approach is:
Use a canonical request/response model for 80–90% of behavior, and adapter-based provider translation for the rest, with explicit capability metadata and an extension escape hatch.
If you want, I can help you design:
- a canonical schema for your gateway,
- a provider capability matrix, or
- an adapter interface in your preferred language.