Prompt
How do I keep vendor lock-in low when using multiple LLM providers?
Latest observation
To keep vendor lock-in low with multiple LLM providers, design your app so the model is a swappable dependency, not the center of your architecture.
Key principles
1. Use a provider-agnostic abstraction
Create an internal interface like:
generate_text(prompt, options)embed(texts)chat(messages)moderate(text)
Then implement adapters for OpenAI, Anthropic, Google, Azure, local models, etc.
Your application should talk to your interface, not the vendor SDK directly.
2. Avoid provider-specific features in core logic
Try not to rely heavily on:
- proprietary prompt formats
- vendor-specific tool/function calling semantics
- special system prompt behaviors
- custom response objects
- vendor-only moderation or memory features
If you use them, isolate them behind adapters and have fallbacks.
3. Standardize your internal request/response schema
Normalize:
- chat messages
- tool calls
- streaming tokens
- usage metrics
- finish reasons
- errors and retries
That way switching vendors doesn’t require rewriting the app.
4. Keep prompts portable
Write prompts so they work across models:
- avoid depending on exact phrasing quirks of one model
- test across vendors
- keep instructions short and explicit
- separate policy/context from model-specific formatting
5. Store raw data, not vendor-derived artifacts
Keep:
- source documents
- user messages
- tool results
- your own embeddings if possible
- evaluation outputs
Don’t make your system depend on opaque vendor-generated memory, cached context, or proprietary indexes unless you can recreate them elsewhere.
6. Abstract model configuration
Use your own config layer for:
- model name aliases
- temperature
- max tokens
- top-p
- timeout
- retry policy
- region/data residency preferences
Example:
fast-chathigh-quality-chatcheap-embedder
Then map those aliases to different providers.
7. Build routing and fallback logic yourself
Have your app decide:
- primary provider
- fallback provider
- per-task routing
- cost/latency-based selection
This reduces dependence on one vendor and gives you portability.
8. Use an evaluation suite
Before committing to a provider, create tests for:
- answer quality
- tool use
- structured output validity
- latency
- refusal behavior
- hallucination rate
If you measure with your own evals, you can switch vendors with less risk.
9. Separate business logic from orchestration
Keep your app’s rules, workflows, and state machines independent of model behavior.
The model should assist, not own the process.
10. Be cautious with fine-tuning and proprietary memory
Fine-tuning can increase lock-in if:
- the training format is vendor-specific
- the deployment path is exclusive
- the dataset is not portable
If you fine-tune, keep datasets, labels, and training scripts portable too.
Practical architecture pattern
A good setup is:
- LLM gateway layer: your own API wrapper
- provider adapters: OpenAI/Anthropic/etc.
- normalized internal schema
- routing layer: picks model based on task
- evaluation layer: compares providers
- observability: logs prompts, outputs, cost, latency
What to avoid
- calling vendor SDKs directly from every service
- hardcoding model IDs in business code
- depending on one provider’s special tool-calling format
- storing vendor-specific conversation state you can’t reconstruct
- using only one provider’s embedding/vector stack if portability matters
A simple rule of thumb
If you can replace Provider A with Provider B by changing only:
- configuration,
- adapter code,
- and maybe a small prompt tweak,
then your lock-in is low.
If you want, I can also give you:
- a reference architecture
- a Python adapter example
- or a checklist for evaluating LLM portability.