Prompt

How do I deploy embeddings, reranking, and retrieval as one stack?

Artificial Intelligence / AI Infrastructure1 observationLast seen Oct 9, 2026

Latest observation

Oct 9, 2026OpenAI APIWeb search: off

A practical way to deploy embeddings + retrieval + reranking as one stack is to treat it as a 3-stage retrieval pipeline behind a single API:

  1. Embed query/documents
  2. Retrieve candidates fast
  3. Rerank the shortlist
  4. Return top results to the app or LLM

Reference architecture

1) Ingestion path

When you index content:

  • chunk documents
  • generate embeddings for each chunk
  • store:
    • vector embedding
    • chunk text
    • metadata
    • document id / tenant id / access control fields

Common storage:

  • Vector DB: Pinecone, Weaviate, Milvus, Qdrant, pgvector
  • Object store / DB for raw text and metadata: Postgres, S3, etc.

2) Query path

When a user searches:

  • embed the query
  • do vector search for top K candidates
  • optionally combine with keyword/BM25 search
  • rerank the top K with a cross-encoder or reranker model
  • return top N

Typical values:

  • retrieve K = 50–200
  • rerank N = 5–20

Deployment patterns

Pattern A: One service, one API

Put everything behind one backend service.

Components

  • API server: FastAPI / Flask / Node
  • Embedding model server: Hugging Face TEI, Sentence Transformers, OpenAI API, vLLM for embedding models
  • Vector database
  • Reranker server: cross-encoder model, Cohere Rerank, Jina reranker, bge-reranker, etc.

Flow

Client -> API -> embed query -> vector DB retrieve -> reranker -> response

This is the simplest production setup.


Pattern B: Separate microservices

Use separate services if you need scale or isolation.

Services

  • embedding-service
  • retrieval-service
  • rerank-service
  • orchestrator-api

Why

  • embeddings scale differently from reranking
  • reranking is often more GPU-heavy
  • retrieval is usually CPU + DB bound

This is the most flexible option for larger systems.


Recommended stack choices

If you want simplest production

  • API: FastAPI
  • Embeddings: SentenceTransformers or managed embedding API
  • Vector DB: Qdrant or pgvector
  • Reranker: bge-reranker or Cohere rerank
  • Orchestration: a single retrieval endpoint

If you want maximum throughput

  • API gateway: FastAPI / Kong / Envoy
  • Embedding model server: TEI or vLLM
  • Vector DB: Milvus / Qdrant / Pinecone
  • Reranker: separate GPU inference service
  • Cache: Redis for repeated queries

How the request flow works

Ingestion

raw docs
  -> chunking
  -> embedding
  -> store vectors + metadata

Query-time

user query
  -> query embedding
  -> vector search top 100
  -> rerank top 100
  -> return top 10

Example implementation outline

1. Embed and store documents

Pseudo-code:

chunks = split_text(document)

embeddings = embed_model.encode(chunks)

for chunk, vector in zip(chunks, embeddings):
    vector_db.upsert({
        "id": chunk.id,
        "vector": vector,
        "payload": {
            "text": chunk.text,
            "doc_id": document.id,
            "source": document.source
        }
    })

2. Query and rerank

query_vec = embed_model.encode([query])[0]

candidates = vector_db.search(
    vector=query_vec,
    top_k=100,
    filter={"tenant_id": tenant_id}
)

reranked = reranker.rank(
    query=query,
    documents=[c["text"] for c in candidates]
)

top_results = reranked[:10]

Important production concerns

1) Latency

Reranking adds latency, so only rerank the shortlist.

Typical:

  • embedding: 5–50 ms
  • vector search: 10–50 ms
  • reranking: 50–300+ ms depending on model and batch size

2) Filtering

Add metadata filters before or during retrieval:

  • tenant
  • language
  • doc type
  • ACL / permissions

3) Hybrid retrieval

Best results often come from:

  • dense vector search
  • BM25 keyword search
  • reranking the union of both

4) Caching

Cache:

  • query embeddings
  • frequent retrieval results
  • reranker results for repeated queries

5) Evaluation

Measure:

  • recall@K
  • MRR
  • nDCG
  • latency p95/p99

A good production blueprint

Minimal stack

  • FastAPI orchestrator
  • Qdrant for vector retrieval
  • SentenceTransformers for embeddings
  • bge-reranker for reranking
  • Redis for caching
  • Postgres for metadata and access control

Request lifecycle

  1. User sends query
  2. API checks permissions
  3. Query embedding generated
  4. Vector DB retrieves top 100
  5. Optional BM25 retrieves top 100
  6. Merge candidates
  7. Rerank top 100
  8. Return top 10 with snippets

If you want to deploy on Kubernetes

Run each part as a deployment:

  • api
  • embedder
  • reranker
  • qdrant / milvus
  • redis
  • postgres

Use:

  • HPA for autoscaling
  • GPU nodes for reranker/embeddings if needed
  • readiness/liveness probes
  • Prometheus + Grafana for observability

Rule of thumb

If you’re starting out:

  • Use one API service
  • Use one vector DB
  • Use one embedding model
  • Use one reranker
  • Add microservices only when scaling requires it

If you want, I can also give you:

  1. a reference architecture diagram
  2. a Docker Compose example
  3. a FastAPI code sample
  4. a Kubernetes deployment template

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.