Prompt
I'm building a RAG product and need advice on the infrastructure stack
Latest observation
A good RAG infrastructure stack is usually built around four layers:
- Ingestion / processing
- Retrieval / storage
- Generation / orchestration
- Evaluation / observability
The best stack depends on your scale, latency target, budget, and whether you’re optimizing for fast MVP or production reliability. Here’s a practical breakdown.
1) Start with the core architecture
A typical RAG setup looks like:
- Source systems: PDFs, docs, wikis, tickets, databases, web pages
- Ingestion pipeline: extract → clean → chunk → enrich metadata → embed
- Indexing layer: vector DB + optional keyword/BM25 index
- Retrieval layer: hybrid search + reranking
- LLM layer: prompt assembly + answer generation
- Monitoring: tracing, quality metrics, cost tracking, feedback loops
2) Recommended stack by layer
A. Ingestion / ETL
Use this for document parsing, chunking, scheduling, and sync.
Common choices
- Python for pipeline logic
- Unstructured or Apache Tika for document parsing
- Airbyte or Fivetran for connectors if you need many sources
- Dagster or Prefect for workflow orchestration
- Celery / RQ / Temporal if you need asynchronous jobs
When to choose what
- Fast MVP: Python + Unstructured + simple cron jobs
- Production pipelines: Dagster/Prefect + queue + retries + idempotency
- Enterprise data integration: Airbyte/Fivetran
B. Embeddings
This affects retrieval quality a lot.
Options
- Managed APIs: OpenAI, Cohere, Voyage, Gemini
- Self-hosted: bge, e5, jina embeddings
Advice
- If you’re early-stage, use a strong managed embedding API first.
- If cost/privacy becomes important, consider self-hosting later.
- Pick one embedding model and benchmark before changing often.
C. Vector store / retrieval storage
This is one of the most important decisions.
Good managed options
- Pinecone: easy, scalable, production-friendly
- Weaviate Cloud
- Qdrant Cloud
- MongoDB Atlas Vector Search
- Elastic Cloud if you need hybrid search and enterprise search features
Self-hosted options
- Qdrant
- Weaviate
- Milvus
- Postgres + pgvector for simpler workloads
Guidance
- MVP / small-medium scale: Postgres + pgvector is often enough.
- Need hybrid search + simple ops: Elasticsearch/OpenSearch or Weaviate/Qdrant.
- High scale, low ops burden: Pinecone or managed Qdrant/Weaviate.
- If your product is already Postgres-centric: pgvector is a very practical choice.
D. Keyword / hybrid retrieval
Pure vector search often isn’t enough.
Add:
- BM25 / lexical search
- Hybrid retrieval
- Metadata filters
- Reranking
Good options
- Elasticsearch/OpenSearch
- Postgres full-text search if small scale
- Vector DBs that support hybrid retrieval
- Separate reranker model: Cohere Rerank, Voyage rerank, bge-reranker, etc.
Best practice
- Use hybrid search for most real products.
- Add a reranker before sending context to the LLM.
- Use metadata filtering aggressively: tenant, source, date, ACL, doc type.
E. LLM / generation layer
This is where you answer the user.
Common choices
- OpenAI
- Anthropic
- Google Gemini
- Azure OpenAI if you need enterprise procurement/security
Orchestration frameworks
- LlamaIndex: great for RAG-heavy apps
- LangChain: flexible, large ecosystem
- Haystack: solid retrieval-focused framework
- Custom orchestration: often best once you stabilize requirements
Advice
- Use a framework to prototype quickly, but don’t let it become a black box.
- Keep your prompt assembly and retrieval logic observable and testable.
F. Caching
Very important for latency and cost.
Use:
- Redis for:
- query result caching
- embedding cache
- session/state cache
- rate limiting
- Optional CDN/cache if you serve static assets or files
G. Metadata, app DB, and ACLs
You’ll almost always need a relational DB.
Use:
- Postgres for:
- users
- tenants
- document metadata
- access control lists
- ingestion state
- audit logs
- evaluation records
Postgres often becomes the system of record, while vector DB handles similarity search.
H. Eventing / queue
For scale and reliability:
- SQS / SNS
- Kafka
- RabbitMQ
- Redis queues for simpler setups
Use queues for:
- ingestion
- reindexing
- embedding generation
- background enrichment
- offline evaluations
I. Observability / evaluation
This is crucial in RAG.
Track:
- retrieval precision/recall
- reranker quality
- groundedness
- citation correctness
- hallucination rate
- latency breakdown
- token/cost per query
- user feedback
Tools:
- OpenTelemetry
- LangSmith
- Arize Phoenix
- Helicone
- PromptLayer
- Datadog / Grafana / Prometheus
- custom eval harness in Python
3) My practical recommendations by stage
If you’re building an MVP
Use:
- Backend: Python + FastAPI
- Orchestration: LlamaIndex or LangChain
- DB: Postgres
- Vector store: pgvector or managed Qdrant/Pinecone
- Embedding: managed API
- LLM: OpenAI/Anthropic
- Queue: Redis/RQ or Celery
- Observability: OpenTelemetry + basic logging
This gets you moving quickly without overengineering.
If you’re building a production SaaS
Use:
- Backend: FastAPI / Node.js / Go depending on team
- DB: Postgres
- Vector store: Pinecone/Qdrant/Weaviate/Elastic depending on retrieval needs
- Search: hybrid retrieval + reranker
- Queue: SQS/Kafka
- Orchestration: Dagster/Temporal/Prefect for ingestion
- Cache: Redis
- Monitoring: OpenTelemetry + Datadog + Phoenix/LangSmith
- Auth/ACL: enforced at retrieval time
Key: make retrieval secure, observable, and testable.
If you’re building for enterprise search
Prioritize:
- ACL-aware retrieval
- auditability
- data residency
- encryption
- hybrid search
- source citations
- incremental indexing
- admin tooling
- evaluation and approval workflows
Likely stack:
- Elastic/OpenSearch or Weaviate/Qdrant + BM25
- Postgres
- Temporal/Dagster
- Azure OpenAI / Anthropic / approved vendor
- Datadog + SIEM integration
- Private networking / VPC deployment
4) The most common mistakes
- Using only vector search and ignoring keyword search
- Not reranking
- Ignoring ACLs and tenant isolation
- No evaluation harness
- No tracing of retrieval and prompts
- Over-chunking or under-chunking documents
- Re-embedding too often without versioning
- Building around framework defaults without understanding them
- Putting all metadata only in the vector DB
- No strategy for stale data or incremental updates
5) A strong default stack
If you want a “safe default” stack that works for many products:
- API: FastAPI
- Workflow: Dagster or Prefect
- DB: Postgres
- Vector: Qdrant or pgvector
- Search: hybrid retrieval + reranker
- Embeddings: managed embedding API
- LLM: OpenAI or Anthropic
- Cache: Redis
- Queue: SQS or Celery
- Observability: OpenTelemetry + Phoenix/LangSmith
- Infra: Docker + Kubernetes or simpler container hosting early on
6) If you want the shortest possible answer
For most RAG products:
- Postgres + vector DB + Redis + queue
- hybrid retrieval + reranking
- managed embeddings and LLMs initially
- strong observability and evals
- ACLs in the retrieval layer
That combination is usually the right starting point.
If you want, I can also give you:
- a cheap MVP stack
- a scalable production stack
- a multi-tenant enterprise stack
- a reference architecture diagram
- a make-vs-buy comparison of vector DBs
Brands