Prompt
I'm trying to build a production copilot over internal data. What architecture works best for freshness, citations, and scale?
Latest observation
For a production copilot over internal data, the best architecture is usually a hybrid retrieval + generation system with separate pipelines for structured and unstructured data, plus strong document provenance and incremental indexing.
Recommended architecture
1) Ingestion layer
Collect data from:
- Docs: Google Drive, Confluence, Notion, SharePoint, Slack exports, PDFs
- Code and tickets: GitHub, Jira, Linear
- Structured sources: SQL databases, warehouse tables, CRM, product telemetry
Normalize into a common format:
document_idsourceauthorcreated_atupdated_atpermissionscontentsection/chunk_idurl/citation target
This metadata is critical for citations, filtering, and freshness.
2) Freshness-first indexing strategy
Use incremental indexing, not full rebuilds:
- Event-driven updates where possible
- CDC for databases
- Webhooks for doc systems
- Scheduled diffs for systems without events
- Tombstone deleted/privileged-removed content immediately
Keep two indexes:
- Hot index for recently changed content, indexed quickly
- Main index for the full corpus
This helps freshness without destabilizing search quality.
If your data is highly time-sensitive, add:
- A cache of recent answers
- A recency-aware reranker
- Query-time lookup into live systems for the latest state
3) Retrieval layer: hybrid search
Use hybrid retrieval:
- Keyword / BM25 for exact terms, IDs, error codes, names
- Vector search for semantic matches
- Optional metadata filters for team, date, doc type, access control
Then rerank results with a stronger model:
- Cross-encoder reranker or LLM-based reranker
- Prioritize:
- relevance
- recency
- authority/source quality
- answerability
- user permissions
This is usually much better than pure vector search.
4) Grounded generation
The LLM should answer only from retrieved sources when possible:
- Pass top passages + metadata into the prompt
- Ask the model to cite each claim
- Require “I don’t know” when evidence is insufficient
- Optionally separate:
- extractive mode for exact citations
- synthesis mode for summaries
For high-stakes use cases, use a verification step:
- Check that each sentence maps to evidence
- Reject unsupported claims
- Format citations per sentence or per paragraph
5) Citation strategy
For reliable citations, store and return:
- Document title
- URL
- Section heading
- Chunk text span offsets if possible
- Timestamp/version
Best practice:
- Cite at the chunk or span level
- Preserve source snippets in the UI
- Show “last updated” and source type
- Prefer citations from the original system of record, not a copied mirror
If possible, use stable anchors:
- paragraph IDs
- section IDs
- line numbers for code
- row IDs for tables
6) Structured data path
Don’t force SQL/data warehouse content into embeddings only.
For structured questions:
- Route to SQL tool execution
- Have the LLM generate safe queries
- Validate queries before execution
- Return result tables + citations to tables/views
- Optionally use semantic layer metrics definitions
This is much more accurate for:
- KPIs
- counts
- time series
- business metrics
- inventory / operational status
A common pattern is:
- Router
- if factual text question → retrieval
- if metric/query question → SQL/tool
- if mixed → both
7) Permissioning and security
Use document-level and row-level access control at retrieval time:
- Filter before the model sees the content
- Never rely on the model to hide restricted data
- Maintain ACL-aware indexes or ACL-aware filtering layers
Also:
- redact secrets/PII where necessary
- log access for auditability
- support tenant/team isolation if needed
8) Scaling pattern
To scale:
- Decouple ingestion, indexing, retrieval, and generation
- Use async queues for ingestion
- Shard indexes by tenant/source if needed
- Cache embeddings and frequent queries
- Keep chunk sizes consistent
- Use a reranker only on a small candidate set
A typical flow:
- User asks question
- Query router classifies intent
- Retriever pulls top candidates from hybrid index
- Reranker orders them
- LLM answers with citations
- Optional verifier checks grounding
- Response returned with source links/snippets
Best-practice reference architecture
Recommended production stack:
- Connectors/CDC → Kafka/SQS/PubSub
- Normalization service
- Chunking + metadata enrichment
- Hybrid index
- BM25/OpenSearch/Elasticsearch
- Vector DB or vector-enabled search engine
- Reranker
- Query router
- LLM generation service
- Citation/verification service
- Permission filter
- Observability + evals
If you want the best tradeoff
For most internal copilots, the sweet spot is:
- Hybrid retrieval
- Incremental fresh indexing
- Tool-based SQL for structured data
- Sentence-level citations
- ACL-aware retrieval
- Verifier for unsupported claims
That gives the best balance of:
- freshness
- traceable citations
- accuracy
- scale
What to avoid
- Pure vector search for everything
- Re-indexing the entire corpus on every change
- Letting the model answer without retrieval
- Ignoring ACLs until after generation
- Converting structured tables into text only
- Over-chunking or under-chunking everything the same way
A practical default architecture
If you want a strong baseline:
- OpenSearch/Elasticsearch for hybrid keyword + vector retrieval
- CDC/webhook ingestion
- Reranker
- LLM with retrieval-only grounding
- SQL tool for metrics
- Citations with URL + section + timestamp
- ACL filtering before retrieval
- Automated evals for freshness and groundedness
If you want, I can also sketch:
- a reference architecture diagram,
- a technology stack recommendation, or
- a step-by-step MVP → production roadmap.