Prompt

How do I build RAG over PDFs, wikis, and tickets with citations?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

To build RAG (Retrieval-Augmented Generation) over PDFs, wikis, and tickets with citations, think of it as 5 layers:

  1. Ingest data
  2. Chunk + enrich
  3. Index for retrieval
  4. Retrieve + rerank
  5. Generate with grounded citations

Below is a practical blueprint.


1) Ingest your sources

You’ll likely have 3 source types:

PDFs

  • Extract text with layout-aware parsers:
    • PyMuPDF, pdfplumber, Unstructured, Apache Tika
  • If PDFs are scanned, run OCR:
    • Tesseract, AWS Textract, Google Document AI, Azure Form Recognizer

Wikis

  • Pull from API or export:
    • Confluence, Notion, MediaWiki, SharePoint, GitHub wiki
  • Preserve:
    • page title
    • section headings
    • page URL
    • last updated time
    • access permissions

Tickets

  • Pull from Jira, Zendesk, ServiceNow, Linear, etc.
  • Preserve:
    • ticket ID
    • title
    • description
    • comments
    • status
    • timestamps
    • assignee/team
    • source link

2) Chunk documents intelligently

Don’t split blindly by fixed length if you want good citations.

Best practice

Chunk by semantic boundaries:

  • PDF: section/subsection/page-aware chunks
  • Wiki: heading-based chunks
  • Tickets: one ticket thread or one issue plus comments, then sub-chunk if too long

Chunk size

A common starting point:

  • 300–800 tokens per chunk
  • Overlap: 50–150 tokens

Each chunk should keep metadata

Store metadata like:

  • source_type: pdf/wiki/ticket
  • doc_id
  • chunk_id
  • title
  • section_heading
  • page_start, page_end
  • url
  • timestamp
  • author
  • permissions

This metadata is what enables citations later.


3) Build your retrieval index

Use a vector store for semantic search, and often a hybrid retriever for better accuracy.

Good options

  • Vector DBs:
    • Pinecone, Weaviate, Qdrant, Milvus, pgvector, Elasticsearch/OpenSearch
  • Embeddings:
    • OpenAI text-embedding models, Cohere, Voyage, bge, e5, etc.

Recommended retrieval strategy

Use hybrid retrieval:

  • BM25 / keyword search for exact terms, IDs, error codes
  • Vector search for semantic matches

Then rerank with a cross-encoder or LLM reranker if needed.

Why hybrid matters

Tickets and wikis often contain exact product names, acronyms, or error codes:

  • Vector search alone may miss them
  • BM25 helps catch them

4) Return citation-ready chunks

When retrieving, you want each result to be a chunk that can be cited directly.

Each retrieved chunk should include:

  • text
  • source title
  • source URL
  • page/section/ticket reference
  • snippet offsets if possible

Example stored record:

{
  "text": "Reset the API key in the Admin Console under Security > API Keys...",
  "metadata": {
    "source_type": "wiki",
    "title": "API Key Management",
    "url": "https://wiki.company.com/api-key-management",
    "section_heading": "Rotating keys",
    "chunk_id": "wiki-123-4"
  }
}

5) Generate answers with grounded citations

Your prompt should instruct the model:

  • Use only retrieved context
  • Cite each claim
  • Say “I don’t know” if context is insufficient

Prompt pattern

Answer the question using only the provided context.
Cite each factual statement with the source ID in square brackets.
If multiple sources support a claim, cite all relevant ones.
If the answer is not in the context, say you couldn’t find it.

Example citation style

  • Inline: The API key expires after 90 days [wiki-123].
  • Sentence-level: one citation per sentence
  • Paragraph-level: cite at the end of each paragraph

A good citation design

Use a source ID per chunk, and map that to human-readable citations.

Example:

  • [pdf:handbook:v3:p12-p13]
  • [wiki:confluence:API-Key-Management#Rotating-keys]
  • [ticket:JIRA-4821]

Then your UI can render:

  • source title
  • page number or section
  • link
  • highlighted passage

Recommended architecture

Ingestion pipeline

  1. Source connectors pull documents
  2. Parse text + structure
  3. Chunk documents
  4. Add metadata
  5. Embed chunks
  6. Store in vector DB + keyword index

Query pipeline

  1. User asks question
  2. Query rewritten/expanded if needed
  3. Hybrid retrieval fetches top-k chunks
  4. Reranker selects best evidence
  5. LLM answers using evidence
  6. Output includes citations and source links

Practical stack examples

Simple and effective

  • Parsing: unstructured, PyMuPDF
  • Embeddings: OpenAI / bge-large
  • Vector DB: Qdrant or pgvector
  • Keyword search: OpenSearch/Elasticsearch
  • RAG framework: LlamaIndex or LangChain
  • Reranking: Cohere rerank / bge-reranker / cross-encoder
  • API/UI: FastAPI + React

Enterprise-ish

  • Source connectors: custom or vendor APIs
  • ETL: Airflow, Dagster
  • Storage: object store for raw docs + Postgres for metadata
  • Search: OpenSearch hybrid
  • Vector: Milvus/Qdrant
  • Authz: ACL filtering before retrieval

Handling permissions

This is important for tickets and internal wikis.

Before retrieval:

  • Filter chunks by user permissions
  • Enforce doc-level and chunk-level ACLs
  • Never let the model see unauthorized content

Recommended:

  • Store ACL metadata on every chunk
  • Apply permission filters in the retriever, not after generation

Common pitfalls

1. Bad chunking

Too large:

  • retrieval becomes noisy
    Too small:
  • loses context

2. Missing metadata

Without page/section/source IDs, citations become weak.

3. Using only vector search

Bad for:

  • error codes
  • ticket IDs
  • exact names
  • config values

4. No reranking

Top-k retrieval often contains near-matches; reranking improves answer quality.

5. Hallucinated citations

Only cite retrieved chunks and verify the cited text actually supports the claim.


How to make citations trustworthy

To ensure the model doesn’t invent citations:

  • Pass source IDs in the context
  • Ask for citations per sentence
  • Post-check that cited source IDs exist in retrieved context
  • Optionally verify claim-to-evidence with an entailment step

A robust approach:

  1. Retrieve evidence
  2. Have the model draft answer with citations
  3. Validate that each citation is among retrieved chunks
  4. Optionally compare answer sentences to cited passages

Minimal implementation outline

Step 1: Ingest and chunk

docs = load_sources()
chunks = chunk_documents(docs)

Step 2: Embed and index

vectors = embed([c.text for c in chunks])
vector_db.upsert(chunks, vectors)

Step 3: Retrieve

results = hybrid_search(query, k=10)
top_chunks = rerank(query, results)

Step 4: Generate with citations

prompt = build_prompt(query, top_chunks)
answer = llm.generate(prompt)

Step 5: Render citations

display(answer, sources=top_chunks)

Example answer format

You can reset the API key from the Admin Console under Security > API Keys [wiki:API-Key-Management#Rotating-keys].
For older accounts, support must rotate the key manually [ticket:JIRA-4821].
The process takes about 10 minutes and invalidates the previous key immediately [pdf:Security-Runbook:p12].


If you want a strong baseline

Start with:

  • hybrid retrieval
  • chunk-level metadata
  • reranking
  • sentence-level citations
  • ACL filtering

That gets you most of the way to a production-quality RAG system.

If you want, I can also give you:

  1. a reference architecture diagram,
  2. a LangChain/LlamaIndex implementation example, or
  3. a production-ready schema for documents, chunks, and citations.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.