Prompt
I'm building a knowledge base with semantic search and citations
Latest observation
Great use case. A knowledge base with semantic search and citations usually works best as a retrieval-augmented system: you index your content for search, retrieve the most relevant passages at query time, and generate answers grounded in those passages with source citations.
Here’s a practical blueprint.
Core architecture
-
Ingest documents
- PDFs, HTML, docs, tickets, wiki pages, etc.
- Extract text plus metadata:
- title
- source URL / file path
- author
- timestamp
- section headings
- page numbers / paragraph IDs
-
Chunk content
- Split docs into semantically meaningful chunks, not just fixed-length text.
- Typical size: 200–800 tokens with overlap.
- Keep chunk-level metadata so citations can point back to the exact source.
-
Create embeddings
- Convert each chunk into a vector embedding.
- Store in a vector index for semantic search.
-
Retrieve
- At query time, embed the user query.
- Search the vector store for top-k relevant chunks.
- Optionally combine with keyword search for hybrid retrieval.
-
Answer with citations
- Feed retrieved chunks to the LLM.
- Require the model to cite only from retrieved sources.
- Return answer plus source references.
Citation strategy
Good citations depend on preserving source structure.
For documents
Store metadata like:
doc_idtitleurlpage_numbersectionchunk_idtext
Then cite in a human-readable form such as:
[1] Employee Handbook, p. 14, “Leave Policy”[2] https://example.com/docs/abc#section-3
For web content
Preserve:
- canonical URL
- page title
- heading anchors
- scrape timestamp
For internal records
Use stable IDs:
- ticket number
- incident ID
- wiki page version
- database record ID
Recommended retrieval flow
A strong production setup often uses:
- Hybrid search
- dense vectors for semantic matching
- sparse/keyword search for exact terms
- Reranking
- re-rank top 20–50 results with a cross-encoder or LLM-based scorer
- Context assembly
- include only the most relevant chunks
- deduplicate overlapping chunks
- keep citation metadata attached
Prompting for grounded answers
Use a strict instruction like:
Answer only using the provided sources.
If the sources do not contain enough information, say so.
Cite every factual claim with source numbers.
This reduces hallucinations and makes citations more reliable.
Good data model
A simple chunk schema:
{
"chunk_id": "doc123_p4_c2",
"doc_id": "doc123",
"title": "Security Policy",
"source": "https://intranet.example.com/security-policy",
"page": 4,
"section": "Access Control",
"text": "Users must change passwords every 90 days...",
"embedding": [ ... ]
}
And query results:
{
"answer": "Users must change passwords every 90 days [1].",
"citations": [
{
"ref": 1,
"title": "Security Policy",
"source": "https://intranet.example.com/security-policy",
"page": 4,
"section": "Access Control",
"chunk_id": "doc123_p4_c2"
}
]
}
Common pitfalls
- Chunks too large: retrieval becomes noisy.
- Chunks too small: citations lose context.
- No stable metadata: citations become unusable.
- Relying only on embeddings: exact terms may be missed.
- Letting the model invent citations: always bind citations to retrieved chunks.
- No versioning: answers may cite outdated content.
Evaluation ideas
Track:
- retrieval precision/recall
- citation correctness
- answer faithfulness
- “not enough information” accuracy
Use test queries with known source passages and verify the system cites the right chunk.
If you want a minimal stack
A common stack:
- Ingestion: Python + BeautifulSoup / PDF parsers
- Embeddings: OpenAI embeddings or similar
- Vector DB: FAISS, pgvector, Pinecone, Weaviate, Milvus
- Reranker: cross-encoder or LLM rerank
- Generation: LLM with strict source-grounding prompt
If you want, I can help with any of these next:
- a reference architecture diagram
- a database schema for citations
- a sample retrieval + citation prompt
- a Python implementation outline