Prompt
How do I deploy embeddings, reranking, and retrieval as one stack?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
A practical way to deploy embeddings + retrieval + reranking as one stack is to treat it as a 3-stage retrieval pipeline behind a single API:
- Embed query/documents
- Retrieve candidates fast
- Rerank the shortlist
- Return top results to the app or LLM
Reference architecture
1) Ingestion path
When you index content:
- chunk documents
- generate embeddings for each chunk
- store:
- vector embedding
- chunk text
- metadata
- document id / tenant id / access control fields
Common storage:
- Vector DB: Pinecone, Weaviate, Milvus, Qdrant, pgvector
- Object store / DB for raw text and metadata: Postgres, S3, etc.
2) Query path
When a user searches:
- embed the query
- do vector search for top
Kcandidates - optionally combine with keyword/BM25 search
- rerank the top
Kwith a cross-encoder or reranker model - return top
N
Typical values:
- retrieve
K = 50–200 - rerank
N = 5–20
Deployment patterns
Pattern A: One service, one API
Put everything behind one backend service.
Components
- API server: FastAPI / Flask / Node
- Embedding model server: Hugging Face TEI, Sentence Transformers, OpenAI API, vLLM for embedding models
- Vector database
- Reranker server: cross-encoder model, Cohere Rerank, Jina reranker, bge-reranker, etc.
Flow
Client -> API -> embed query -> vector DB retrieve -> reranker -> response
This is the simplest production setup.
Pattern B: Separate microservices
Use separate services if you need scale or isolation.
Services
embedding-serviceretrieval-servicererank-serviceorchestrator-api
Why
- embeddings scale differently from reranking
- reranking is often more GPU-heavy
- retrieval is usually CPU + DB bound
This is the most flexible option for larger systems.
Recommended stack choices
If you want simplest production
- API: FastAPI
- Embeddings: SentenceTransformers or managed embedding API
- Vector DB: Qdrant or pgvector
- Reranker: bge-reranker or Cohere rerank
- Orchestration: a single retrieval endpoint
If you want maximum throughput
- API gateway: FastAPI / Kong / Envoy
- Embedding model server: TEI or vLLM
- Vector DB: Milvus / Qdrant / Pinecone
- Reranker: separate GPU inference service
- Cache: Redis for repeated queries
How the request flow works
Ingestion
raw docs
-> chunking
-> embedding
-> store vectors + metadata
Query-time
user query
-> query embedding
-> vector search top 100
-> rerank top 100
-> return top 10
Example implementation outline
1. Embed and store documents
Pseudo-code:
chunks = split_text(document)
embeddings = embed_model.encode(chunks)
for chunk, vector in zip(chunks, embeddings):
vector_db.upsert({
"id": chunk.id,
"vector": vector,
"payload": {
"text": chunk.text,
"doc_id": document.id,
"source": document.source
}
})
2. Query and rerank
query_vec = embed_model.encode([query])[0]
candidates = vector_db.search(
vector=query_vec,
top_k=100,
filter={"tenant_id": tenant_id}
)
reranked = reranker.rank(
query=query,
documents=[c["text"] for c in candidates]
)
top_results = reranked[:10]
Important production concerns
1) Latency
Reranking adds latency, so only rerank the shortlist.
Typical:
- embedding: 5–50 ms
- vector search: 10–50 ms
- reranking: 50–300+ ms depending on model and batch size
2) Filtering
Add metadata filters before or during retrieval:
- tenant
- language
- doc type
- ACL / permissions
3) Hybrid retrieval
Best results often come from:
- dense vector search
- BM25 keyword search
- reranking the union of both
4) Caching
Cache:
- query embeddings
- frequent retrieval results
- reranker results for repeated queries
5) Evaluation
Measure:
- recall@K
- MRR
- nDCG
- latency p95/p99
A good production blueprint
Minimal stack
- FastAPI orchestrator
- Qdrant for vector retrieval
- SentenceTransformers for embeddings
- bge-reranker for reranking
- Redis for caching
- Postgres for metadata and access control
Request lifecycle
- User sends query
- API checks permissions
- Query embedding generated
- Vector DB retrieves top 100
- Optional BM25 retrieves top 100
- Merge candidates
- Rerank top 100
- Return top 10 with snippets
If you want to deploy on Kubernetes
Run each part as a deployment:
apiembedderrerankerqdrant/milvusredispostgres
Use:
- HPA for autoscaling
- GPU nodes for reranker/embeddings if needed
- readiness/liveness probes
- Prometheus + Grafana for observability
Rule of thumb
If you’re starting out:
- Use one API service
- Use one vector DB
- Use one embedding model
- Use one reranker
- Add microservices only when scaling requires it
If you want, I can also give you:
- a reference architecture diagram
- a Docker Compose example
- a FastAPI code sample
- a Kubernetes deployment template