Prompt
I'm building a multi-source retrieval pipeline; recommend the right stack
Latest observation
Here’s a practical stack recommendation for a multi-source retrieval pipeline. I’ll optimize for relevance, scalability, and maintainability rather than “most trendy.”
1) Core architecture I’d recommend
Use a hybrid retrieval stack:
- Ingestion/connectors: pull from APIs, docs, DBs, object stores, web, etc.
- Normalization layer: convert everything into a canonical document schema.
- Chunking + metadata enrichment: split content into retrievable units with source/time/permission metadata.
- Indexing:
- BM25 / keyword search for exact matches
- Vector embeddings for semantic search
- optionally reranking for final ordering
- Fusion / retrieval orchestration: combine results from multiple sources and modalities.
- Governance layer: deduplication, freshness, ACLs, provenance, auditability.
This pattern usually beats “vector-only.”
2) Recommended stack by layer
A. Ingestion / connectors
Pick based on source diversity:
- Airbyte or Fivetran for standard SaaS/DB ingestion
- Custom Python connectors for APIs, files, and web sources
- Dagster or Airflow for orchestration
Recommendation:
- If you want flexibility and engineering control: Dagster + custom connectors
- If you want managed data movement for common sources: Airbyte
B. Document processing / normalization
- Python
- Pydantic for schema validation
- Unstructured for PDFs, HTML, DOCX, email, etc.
- Apache Tika if you need broad file parsing
- OCR: Tesseract, AWS Textract, or Google Document AI
Canonical schema should include:
doc_idsourcesource_typetitletextchunksmetadatacreated_atupdated_atacl / permissionsprovenance
C. Chunking and enrichment
- LlamaIndex or LangChain for chunking utilities and document loaders
- Or build your own simple chunker if you want tight control
Best practice:
- chunk by semantic boundaries where possible
- keep chunk sizes moderate
- preserve parent document references
- attach metadata to every chunk
Add enrichments like:
- language detection
- entity extraction
- timestamps
- source confidence
- access scope
D. Indexing / storage
For multi-source retrieval, I’d usually use two indexes:
1. Keyword index
- OpenSearch / Elasticsearch
- Great for exact terms, filters, faceting, auditability
2. Vector index
Options:
- Pinecone: managed, simple
- Weaviate: flexible, good hybrid features
- Milvus: strong open-source choice at scale
- pgvector: great if you want to stay inside Postgres and your scale is moderate
Recommendation by scale:
- Small to medium: Postgres + pgvector + OpenSearch
- Medium to large: OpenSearch + Pinecone/Weaviate/Milvus
- If you want simplest managed setup: Pinecone + OpenSearch
E. Embeddings
Use a strong general-purpose embedding model:
- OpenAI text-embedding-3-large / small
- Cohere embeddings
- BGE-M3 or e5 family if you want open-source/self-hosted
Recommendation:
- Fastest to production: hosted embeddings
- Highest control / lower cost at scale: BGE-M3 or e5-large
F. Retrieval orchestration
Use hybrid retrieval:
- BM25 top-k
- vector top-k
- merge using weighted score fusion
- rerank top candidates with a cross-encoder / reranker
Useful tools:
- Elastic/OpenSearch hybrid search
- RRF (Reciprocal Rank Fusion) for merging results
- Cohere Rerank or open-source rerankers
- LlamaIndex if you want retrieval abstractions quickly
Best practice: rerank after fusion, not before.
G. Reranking
This is often the biggest relevance boost.
Options:
- Cohere Rerank
- bge-reranker
- cross-encoder/ms-marco variants
Use reranking on:
- top 20–100 fused candidates
- with query + chunk text + metadata
H. Serving / API layer
- FastAPI for retrieval service
- Redis for caching
- Postgres for metadata, ACLs, jobs, lineage
- Kafka / RabbitMQ / SQS if you need event-driven ingestion
I. Observability / evaluation
Don’t skip this.
- OpenTelemetry for traces
- Prometheus + Grafana for metrics
- LangSmith or Arize Phoenix for retrieval/debugging evals
- Build offline evaluation with:
- recall@k
- MRR
- nDCG
- answer grounding metrics
3) A good default stack
If you want a strong “just build it” recommendation:
Best balanced stack
- Orchestration: Dagster
- Parsing: Unstructured + custom Python
- Metadata store: Postgres
- Keyword search: OpenSearch
- Vector store: pgvector (or Pinecone if managed)
- Embeddings: OpenAI or BGE-M3
- Reranker: Cohere Rerank or bge-reranker
- API: FastAPI
- Cache: Redis
- Observability: OpenTelemetry + Grafana
- Retrieval framework: custom orchestration, with optional LlamaIndex/LangChain helpers
This gives you:
- good hybrid retrieval
- easy metadata filtering
- source provenance
- scalability without overcomplicating the system
4) If you want fewer moving parts
Lean stack
- Postgres + pgvector
- OpenSearch (or Postgres full-text if smaller)
- FastAPI
- Dagster
- Python + Unstructured
Use this if:
- your corpus is under a few million chunks
- you want operational simplicity
- you can tolerate moderate search sophistication
5) If you want maximum scale
Scalable stack
- Kafka for ingestion events
- Spark / Ray for processing
- OpenSearch for text search
- Milvus / Weaviate / Pinecone for vectors
- Postgres for metadata/ACLs
- Dagster for orchestration
- FastAPI serving layer
Use this if:
- you have many sources
- frequent updates
- high QPS
- multiple teams consuming the pipeline
6) My strongest recommendation on retrieval strategy
For multi-source retrieval, use:
- Source-aware filtering
- filter by freshness, permissions, source trust
- Hybrid search
- BM25 + vectors
- Deduplication
- same document from multiple sources
- Reranking
- to improve final ordering
- Provenance returned with every result
- crucial for debugging and trust
7) Common mistakes to avoid
- Going vector-only
- Ignoring ACLs and provenance
- Chunking too aggressively
- No reranking
- Mixing all sources into one undifferentiated index
- Not separating metadata from content
- No offline evaluation set
- No freshness strategy for updates/deletions
8) If you tell me these 5 things, I can give a much sharper stack
- Source types: SaaS, PDFs, web, DBs, tickets, code, etc.
- Corpus size: docs/chunks and update frequency
- Query patterns: exact lookup vs semantic Q&A vs filtering
- Latency target: sub-second, a few seconds, batch
- Deployment preference: managed cloud vs self-hosted
If you want, I can also give you:
- a reference architecture diagram
- a cloud-native stack on AWS/GCP/Azure
- or a minimal MVP stack vs production stack comparison.