Prompt
How do I build RAG over PDFs, wikis, and tickets with citations?
Latest observation
To build RAG (Retrieval-Augmented Generation) over PDFs, wikis, and tickets with citations, think of it as 5 layers:
- Ingest data
- Chunk + enrich
- Index for retrieval
- Retrieve + rerank
- Generate with grounded citations
Below is a practical blueprint.
1) Ingest your sources
You’ll likely have 3 source types:
PDFs
- Extract text with layout-aware parsers:
PyMuPDF,pdfplumber,Unstructured,Apache Tika
- If PDFs are scanned, run OCR:
Tesseract, AWS Textract, Google Document AI, Azure Form Recognizer
Wikis
- Pull from API or export:
- Confluence, Notion, MediaWiki, SharePoint, GitHub wiki
- Preserve:
- page title
- section headings
- page URL
- last updated time
- access permissions
Tickets
- Pull from Jira, Zendesk, ServiceNow, Linear, etc.
- Preserve:
- ticket ID
- title
- description
- comments
- status
- timestamps
- assignee/team
- source link
2) Chunk documents intelligently
Don’t split blindly by fixed length if you want good citations.
Best practice
Chunk by semantic boundaries:
- PDF: section/subsection/page-aware chunks
- Wiki: heading-based chunks
- Tickets: one ticket thread or one issue plus comments, then sub-chunk if too long
Chunk size
A common starting point:
- 300–800 tokens per chunk
- Overlap: 50–150 tokens
Each chunk should keep metadata
Store metadata like:
source_type: pdf/wiki/ticketdoc_idchunk_idtitlesection_headingpage_start,page_endurltimestampauthorpermissions
This metadata is what enables citations later.
3) Build your retrieval index
Use a vector store for semantic search, and often a hybrid retriever for better accuracy.
Good options
- Vector DBs:
- Pinecone, Weaviate, Qdrant, Milvus, pgvector, Elasticsearch/OpenSearch
- Embeddings:
- OpenAI text-embedding models, Cohere, Voyage, bge, e5, etc.
Recommended retrieval strategy
Use hybrid retrieval:
- BM25 / keyword search for exact terms, IDs, error codes
- Vector search for semantic matches
Then rerank with a cross-encoder or LLM reranker if needed.
Why hybrid matters
Tickets and wikis often contain exact product names, acronyms, or error codes:
- Vector search alone may miss them
- BM25 helps catch them
4) Return citation-ready chunks
When retrieving, you want each result to be a chunk that can be cited directly.
Each retrieved chunk should include:
- text
- source title
- source URL
- page/section/ticket reference
- snippet offsets if possible
Example stored record:
{
"text": "Reset the API key in the Admin Console under Security > API Keys...",
"metadata": {
"source_type": "wiki",
"title": "API Key Management",
"url": "https://wiki.company.com/api-key-management",
"section_heading": "Rotating keys",
"chunk_id": "wiki-123-4"
}
}
5) Generate answers with grounded citations
Your prompt should instruct the model:
- Use only retrieved context
- Cite each claim
- Say “I don’t know” if context is insufficient
Prompt pattern
Answer the question using only the provided context.
Cite each factual statement with the source ID in square brackets.
If multiple sources support a claim, cite all relevant ones.
If the answer is not in the context, say you couldn’t find it.
Example citation style
- Inline:
The API key expires after 90 days [wiki-123]. - Sentence-level: one citation per sentence
- Paragraph-level: cite at the end of each paragraph
A good citation design
Use a source ID per chunk, and map that to human-readable citations.
Example:
[pdf:handbook:v3:p12-p13][wiki:confluence:API-Key-Management#Rotating-keys][ticket:JIRA-4821]
Then your UI can render:
- source title
- page number or section
- link
- highlighted passage
Recommended architecture
Ingestion pipeline
- Source connectors pull documents
- Parse text + structure
- Chunk documents
- Add metadata
- Embed chunks
- Store in vector DB + keyword index
Query pipeline
- User asks question
- Query rewritten/expanded if needed
- Hybrid retrieval fetches top-k chunks
- Reranker selects best evidence
- LLM answers using evidence
- Output includes citations and source links
Practical stack examples
Simple and effective
- Parsing:
unstructured,PyMuPDF - Embeddings: OpenAI / bge-large
- Vector DB: Qdrant or pgvector
- Keyword search: OpenSearch/Elasticsearch
- RAG framework: LlamaIndex or LangChain
- Reranking: Cohere rerank / bge-reranker / cross-encoder
- API/UI: FastAPI + React
Enterprise-ish
- Source connectors: custom or vendor APIs
- ETL: Airflow, Dagster
- Storage: object store for raw docs + Postgres for metadata
- Search: OpenSearch hybrid
- Vector: Milvus/Qdrant
- Authz: ACL filtering before retrieval
Handling permissions
This is important for tickets and internal wikis.
Before retrieval:
- Filter chunks by user permissions
- Enforce doc-level and chunk-level ACLs
- Never let the model see unauthorized content
Recommended:
- Store ACL metadata on every chunk
- Apply permission filters in the retriever, not after generation
Common pitfalls
1. Bad chunking
Too large:
- retrieval becomes noisy
Too small: - loses context
2. Missing metadata
Without page/section/source IDs, citations become weak.
3. Using only vector search
Bad for:
- error codes
- ticket IDs
- exact names
- config values
4. No reranking
Top-k retrieval often contains near-matches; reranking improves answer quality.
5. Hallucinated citations
Only cite retrieved chunks and verify the cited text actually supports the claim.
How to make citations trustworthy
To ensure the model doesn’t invent citations:
- Pass source IDs in the context
- Ask for citations per sentence
- Post-check that cited source IDs exist in retrieved context
- Optionally verify claim-to-evidence with an entailment step
A robust approach:
- Retrieve evidence
- Have the model draft answer with citations
- Validate that each citation is among retrieved chunks
- Optionally compare answer sentences to cited passages
Minimal implementation outline
Step 1: Ingest and chunk
docs = load_sources()
chunks = chunk_documents(docs)
Step 2: Embed and index
vectors = embed([c.text for c in chunks])
vector_db.upsert(chunks, vectors)
Step 3: Retrieve
results = hybrid_search(query, k=10)
top_chunks = rerank(query, results)
Step 4: Generate with citations
prompt = build_prompt(query, top_chunks)
answer = llm.generate(prompt)
Step 5: Render citations
display(answer, sources=top_chunks)
Example answer format
You can reset the API key from the Admin Console under Security > API Keys [wiki:API-Key-Management#Rotating-keys].
For older accounts, support must rotate the key manually [ticket:JIRA-4821].
The process takes about 10 minutes and invalidates the previous key immediately [pdf:Security-Runbook:p12].
If you want a strong baseline
Start with:
- hybrid retrieval
- chunk-level metadata
- reranking
- sentence-level citations
- ACL filtering
That gets you most of the way to a production-quality RAG system.
If you want, I can also give you:
- a reference architecture diagram,
- a LangChain/LlamaIndex implementation example, or
- a production-ready schema for documents, chunks, and citations.
Brands