Prompt
production RAG pipeline
Latest observation
Oct 11, 2026
OpenAI APIWeb search: off
A production RAG (Retrieval-Augmented Generation) pipeline usually has these stages:
-
Ingest
- Pull data from sources: PDFs, docs, web pages, databases, tickets, wikis, code, etc.
- Normalize/clean text, extract metadata, deduplicate.
-
Chunk
- Split documents into retrieval-friendly pieces.
- Commonly 200–800 tokens with overlap, but depends on content type.
- Preserve structure and metadata: source, section, timestamp, permissions.
-
Embed + Index
- Convert chunks into embeddings.
- Store in a vector database or hybrid search index.
- Often use hybrid retrieval: dense vectors + BM25 keyword search.
- Build indexes for metadata filtering and ACLs.
-
Query Understanding
- Detect intent, rewrite the query if needed.
- Apply filters like user permissions, time range, product, language.
- Optionally expand acronyms or normalize entities.
-
Retrieve
- Fetch top-k candidate chunks using hybrid search.
- Rerank with a cross-encoder / LLM reranker.
- Keep citations and provenance.
-
Context Assembly
- Deduplicate results.
- Fit best chunks into the model context window.
- Add instructions, citations, and guardrails.
- Possibly summarize or compress retrieved text.
-
Generate
- Call the LLM with the retrieved context.
- Use grounded prompting: answer only from provided sources when required.
- Return citations, confidence, and “I don’t know” when evidence is insufficient.
-
Post-process
- Format answer, citations, links.
- Apply safety checks, policy filters, and hallucination checks.
- Optional structured output parsing.
-
Observe + Evaluate
- Log retrieval hits, latency, token usage, answer quality, user feedback.
- Run offline evals: recall@k, MRR, faithfulness, answer correctness.
- Monitor drift and stale content.
Typical production architecture
- Data layer: object storage / warehouse / document store
- Processing layer: ETL jobs, parsers, chunkers, embedding workers
- Index layer: vector DB + keyword search + metadata store
- Serving layer: query router, retriever, reranker, LLM gateway
- Governance: ACLs, auditing, PII redaction, encryption
- Ops: caching, rate limits, retries, observability, cost controls
Key production concerns
- Freshness: incremental re-indexing and deletion handling
- Access control: retrieval must respect user/document permissions
- Hybrid retrieval: improves robustness over pure vector search
- Reranking: often critical for quality
- Citation quality: essential for trust and debugging
- Evaluation: measure retrieval and generation separately
- Latency/cost: caching and routing to smaller models when possible
Minimal request flow
User query -> query rewrite -> retrieve top-k -> rerank -> assemble context -> LLM generate -> cite sources -> return answer
If you want, I can also provide:
- a reference architecture diagram,
- a production-ready tech stack, or
- sample code for a RAG pipeline.