Prompt
What should I use for a production retrieval layer?
Latest observation
Sep 28, 2026
OpenAI APIWeb search: off
For a production retrieval layer, the best choice depends on your scale, latency target, and how “semantic” your search needs to be. In most modern production systems, I’d recommend this pattern:
Best default choice
Hybrid retrieval = lexical search + vector search + reranking
Typical stack
- Lexical index: Elasticsearch / OpenSearch / PostgreSQL full-text
- Vector index: pgvector, Pinecone, Weaviate, Milvus, Qdrant, OpenSearch k-NN, or Elasticsearch vectors
- Reranker: cross-encoder / LLM reranker / learned ranking model
This gives you:
- high precision on exact terms, IDs, names, and filters
- semantic matching for paraphrases and vague queries
- better ranking quality than either method alone
What to use in practice
1. If you want the simplest production setup
Postgres + pgvector + full-text search
- Great for small-to-mid scale
- Easy operations if you already use Postgres
- Good enough for many internal apps and MVPs
Use this if you need:
- under a few million chunks/documents
- moderate QPS
- strong operational simplicity
2. If you want robust search at scale
OpenSearch / Elasticsearch + vector search + reranker
- Strong for enterprise retrieval
- Excellent filtering, faceting, and keyword matching
- Mature operational story
- Better for larger catalogs and mixed search workloads
Use this if you need:
- complex metadata filters
- auditing, analytics, and operational search features
- high query throughput
3. If you want managed vector retrieval quickly
Pinecone
- Very easy to productionize
- Good managed scaling and reliability
- Less infra overhead
Use this if:
- you want speed of deployment
- you don’t want to manage indexes yourself
4. If you want open-source vector DB
Qdrant or Milvus
- Qdrant: simpler, very practical, strong filtering
- Milvus: more heavyweight, good for large-scale vector workloads
Use this if:
- you want control and self-hosting
- vector search is the core of your application
My recommendation by scenario
RAG / internal knowledge assistant
- Best default: Postgres + pgvector + full-text + reranker
- At larger scale: OpenSearch + vectors + reranker
Product search / e-commerce
- Best default: Elasticsearch/OpenSearch hybrid search
- Add vector search for semantic matching
- Rerank top candidates
High-scale SaaS with strong uptime requirements
- Managed: Pinecone or Elasticsearch/OpenSearch Cloud
- Add reranking layer
Fast-moving startup
- Start with: Postgres + pgvector
- Migrate to OpenSearch/Pinecone/Qdrant when needed
Important production features to require
Whatever you choose, make sure it supports:
- metadata filtering
- incremental updates
- deletion and TTL
- hybrid query support
- observability / query tracing
- index rebuild strategy
- access control if needed
- reranking support or easy integration
A practical architecture
- Ingest documents
- Chunk and embed them
- Store:
- text + metadata in primary DB
- vectors in vector store
- optionally lexical index for keyword search
- Query:
- retrieve top-k from lexical and vector indexes
- merge results
- rerank
- return final context
Short answer
If you want the best general production retrieval layer: Use hybrid retrieval with OpenSearch/Elasticsearch or Postgres+pgvector, plus a reranker.
If you want, I can give you:
- a specific recommendation based on your scale and use case,
- a comparison table of Pinecone/Qdrant/Milvus/pgvector/OpenSearch,
- or a reference architecture for RAG retrieval in production.