Prompt
I'm building a knowledge assistant for engineers. What's a good architecture for hybrid retrieval and citations?
Latest observation
A solid architecture for an engineering knowledge assistant is usually:
- Retrieve broadly with hybrid search
- Rerank for relevance
- Generate with grounded citations
- Keep provenance at the chunk level
- Use answer-time citation mapping and verification
Here’s a practical blueprint.
1) Ingestion and normalization
Sources
Pull from:
- PDFs, specs, design docs
- wiki pages
- tickets / RFCs / postmortems
- code comments, READMEs, runbooks
- API docs and internal knowledge bases
Normalize everything into a common document model
For each document, store:
doc_idtitlesource_urlauthor,team,created_at,updated_atversionor commit hash- ACL / permissions
- section hierarchy
- raw text
- structured fields when available
Chunk with structure awareness
Chunk by:
- section headers
- paragraphs / subheadings
- table boundaries
- code blocks separately
Keep:
chunk_iddoc_idsection_pathchunk_textoffsetsor anchor infopage/ line references for PDFs and codetimestamp/ version
This is essential for citations later.
2) Build multiple indexes
For engineering use cases, hybrid retrieval should usually mean:
A. Sparse index
Use BM25 or similar keyword retrieval for:
- exact terms
- part numbers
- error codes
- API names
- symbols, constants, acronyms
B. Dense vector index
Use embeddings for:
- semantic matching
- paraphrases
- concept-level retrieval
C. Optional structured index
If your docs are rich enough, also index:
- metadata filters
- tags
- service names
- teams
- product versions
- dates
- ACLs
This lets you constrain search by scope.
3) Retrieval strategy: hybrid + filters + rerank
A good default pipeline:
Step 1: Query understanding
Classify the query:
- troubleshooting
- “how do I”
- policy/compliance
- API lookup
- architecture question
- comparison
- code-centric
Optionally expand query with:
- acronyms
- synonyms
- likely product/version names
- normalized error strings
Step 2: Candidate retrieval
Run both:
- BM25 top N
- Dense top N
Merge results with a weighted strategy:
- Reciprocal Rank Fusion (RRF) is a strong default
- or weighted score blending
Step 3: Apply filters
Filter by:
- access control
- recency
- product version
- subsystem
- language
- document type
Step 4: Rerank
Use a cross-encoder or reranker model over the merged candidates.
This is often the biggest quality win.
Step 5: Diversity and deduplication
Avoid returning 10 chunks from the same page. Prefer:
- one or two chunks per document
- coverage across distinct docs when appropriate
4) Citation-ready grounding layer
The key to good citations is to retrieve and store evidence in a way that can be traced back precisely.
Store provenance per chunk
Each chunk should carry:
- source document
- URL
- section title
- page/line/span offsets
- version hash
- access policy
Canonical citation object
At answer time, every cited statement should map to a citation record like:
{
"citation_id": "c1",
"doc_id": "doc_123",
"chunk_id": "chunk_456",
"source_url": "https://internal/wiki/abc",
"title": "Caching Strategy",
"section": "Eviction Policy",
"page": 12,
"lines": [210, 238],
"version": "gitsha-9f2a1"
}
This allows:
- human-readable citations
- clickable links
- exact traceability
- versioned answers
5) Generation: answer only from retrieved evidence
Use a constrained prompt pattern:
- Show the model the question
- Provide the top evidence chunks
- Instruct it to answer only from those chunks
- Require citations on each claim or paragraph
A reliable pattern is:
- One claim = one citation
- If a sentence depends on two sources, cite both
- If evidence is missing, say so explicitly
This reduces hallucinations and citation drift.
6) Citation placement strategy
There are two common patterns:
Inline citations
Example:
The service uses a 5-minute TTL for session keys [Doc A, §3.2].
Best for:
- precise answers
- technical audiences
- short responses
Paragraph-level citations
Example:
The cache invalidation logic is event-driven and backed by Redis. [1]
Best for:
- smoother reading
- longer summaries
For engineers, inline citations are usually better.
7) Post-generation verification
Don’t trust the model blindly.
Run an answer verification step:
- Check that each cited chunk actually supports the claim
- Detect uncited factual statements
- Flag unsupported or overgeneralized claims
- Optionally force regeneration if grounding is weak
This can be done with:
- rule-based checks
- a second LLM verifier
- entailment / QA consistency checks
8) Recommended architecture pattern
A strong production architecture looks like this:
Offline
- Ingest docs
- Parse and normalize
- Chunk with structure
- Create BM25 index
- Create embeddings index
- Store metadata + ACL + provenance
Online
- User asks question
- Query classifier / normalizer
- Hybrid retrieve (BM25 + vector)
- Merge via RRF
- Filter by ACL/version/metadata
- Rerank top candidates
- Select evidence set
- Generate answer with citations
- Verify citations
- Return answer + source list
9) Practical engineering recommendations
Use BM25 heavily for technical content
Engineering queries often contain exact terms:
- function names
- config keys
- stack traces
- error messages
- standards references
Sparse retrieval is very valuable here.
Use dense retrieval for concept questions
Examples:
- “Why are our deploys flaky?”
- “What’s the right pattern for retries?”
- “How do we avoid duplicate writes?”
Keep chunk sizes moderate
A good starting point:
- 200–500 tokens per chunk
- preserve headings and nearby context
- avoid splitting code blocks or tables
Capture section structure
A citation like:
Document > Section > Subsectionis much more useful than just a chunk ID.
Prefer fresh docs when relevance ties
In engineering environments, stale docs are dangerous.
Enforce ACL before generation
Never retrieve or cite content the user cannot access.
10) Good default retrieval settings
A reasonable starting point:
- BM25 top 50–100
- Dense top 50–100
- Merge with RRF
- Rerank top 20–40
- Feed top 5–10 chunks to the generator
If answers require synthesis across docs, include:
- 1–2 chunks per document
- some diversity constraint
11) Handling citations in the UI
Make citations usable:
- clickable source links
- hover preview with highlighted excerpt
- page/line references
- “why this source?” explanation
- confidence / recency indicators
For engineers, a source panel is very effective:
- Source 1: doc title, section, excerpt
- Source 2: same
- Source 3: same
12) Common failure modes to avoid
Over-relying on embeddings
Dense search alone misses exact match queries.
No provenance at chunk level
If you only store doc-level metadata, citations become vague.
Too-large chunks
The model may cite irrelevant text or mix multiple claims.
No reranking
Hybrid retrieval without reranking often returns noisy evidence.
No verification
The model may produce fluent but unsupported answers.
Stale docs without versioning
Citations become misleading when docs change.
13) A simple reference architecture
If you want a minimal-but-good production design:
- Ingestion: document parser + metadata extractor
- Storage: object store for raw docs, database for metadata/chunks
- Search: BM25 index + vector DB
- Retrieval: hybrid merge + ACL filters
- Rerank: cross-encoder
- Generation: citation-constrained LLM
- Verification: entailment / citation checker
- UI: answer + evidence panel
14) My recommended “best default” for engineers
If you want a single recommendation:
- Use BM25 + dense embeddings
- Merge with RRF
- Apply ACL and metadata filters early
- Rerank with a cross-encoder
- Generate with chunk-level citations
- Verify every claim against retrieved evidence
- Keep versioned provenance for every chunk
That combination usually gives the best balance of:
- recall
- precision
- explainability
- citation quality
If you want, I can also sketch:
- a concrete system diagram,
- a retrieval scoring formula,
- a citation schema,
- or a reference implementation stack in Python.