Prompt
ChatGPT: I'm moving from a prototype to production and need to serve embeddings and an LLM with minimal rewrite. What architecture makes sense?
Latest observation
A good “minimal rewrite” production path is to separate the application from the model serving layer and keep the model APIs as close as possible to what you already use in prototyping.
Recommended architecture
1) App layer
Your app should talk to a single internal API that abstracts:
- embeddings
- LLM generation
- reranking / moderation / tools if needed
This keeps your product code stable even if you swap model providers later.
2) Model serving layer
Use a dedicated inference server for both embeddings and LLMs, ideally behind one gateway.
Common setup:
- API gateway / router
- LLM inference server for generation
- Embedding inference server for vectorization
- Vector DB for retrieval
- Cache for repeated prompts/embeddings
- Job queue for async batch work
3) Storage and retrieval
- Store embeddings in a vector database like pgvector, Milvus, Pinecone, Weaviate, or Qdrant.
- Keep raw documents in object storage or your primary DB.
- Add a retrieval service for RAG that handles chunking, embedding, search, and ranking.
Best way to minimize rewrite
Option A: Use an OpenAI-compatible serving interface
This is usually the easiest migration path.
If your prototype already uses OpenAI-style calls, choose serving components that expose:
/v1/chat/completions/v1/embeddings
Examples:
- vLLM for LLM serving
- TGI (Text Generation Inference) for LLM serving
- text-embeddings-inference (TEI) for embeddings
This lets you keep client code largely unchanged.
Option B: Put an internal “AI gateway” in front
Create one internal service that normalizes requests:
generate()embed()rerank()
Then behind the scenes route to:
- hosted API
- self-hosted GPU service
- fallback model
- different providers by cost or latency
This is the cleanest long-term design.
Reference production architecture
Client/App
↓
AI Gateway / Orchestrator
├── LLM Serving (vLLM/TGI)
├── Embedding Serving (TEI/other)
├── Vector DB (Qdrant/Milvus/pgvector)
├── Cache (Redis)
└── Queue/Workers (Celery/Kafka/SQS)
Practical recommendation by stage
If you want fastest path to production
- Keep your current app code
- Wrap model calls behind a thin internal client
- Use OpenAI-compatible serving for both embeddings and LLMs
- Use a managed vector DB
- Add Redis caching
If you need scale/control
- Self-host the models with vLLM + TEI
- Put an API gateway in front
- Add autoscaling GPU workers
- Use observability from day one
Key production concerns
Make sure your architecture handles:
- latency: batching, caching, streaming
- throughput: async queues, autoscaling, GPU utilization
- fault tolerance: retries, fallbacks, timeouts
- versioning: model/version aliases
- observability: traces, prompt logs, token usage, quality evals
- security: auth, rate limits, data redaction
My default recommendation
If you want minimal rewrite and a sane production path:
- App code talks to an internal AI API
- Use OpenAI-compatible endpoints
- Serve LLM with vLLM
- Serve embeddings with TEI
- Use Qdrant or pgvector for retrieval
- Redis for cache
- A small gateway service for routing and observability
That gives you a clean path from prototype to production without locking you in.
If you want, I can sketch:
- a concrete AWS/GCP/Azure deployment, or
- a sample service diagram with specific components, or
- a migration plan from prototype code to this architecture.