Prompt
I'm building an AI copilot for employees; what retrieval setup makes sense?
Latest observation
For an employee copilot, the retrieval setup should usually be hybrid, permission-aware, and source-grounded rather than “just vector search.”
Recommended default setup
1) Use a hybrid retrieval stack
Combine:
- Keyword search for exact terms, IDs, policy names, error codes, people/project names
- Vector search for semantic matches
- Metadata filters for department, region, doc type, date, confidentiality, etc.
Why: employee questions are often a mix of:
- “Where’s the PTO policy?”
- “What’s the process for reimbursing travel?”
- “Find the doc about Q3 pricing changes”
- “Who owns the Kafka migration?”
A hybrid system is much more reliable than embeddings alone.
2) Retrieve from authoritative enterprise sources first
Prioritize:
- HR policies
- Confluence/Notion/wiki docs
- Google Drive / SharePoint
- Ticketing/incident systems
- Internal docs, SOPs, runbooks
- CRM/ERP systems if relevant
- Slack/Teams only with caution and clear provenance
Best practice:
- Define a source hierarchy
- Mark certain sources as “canonical”
- Prefer recent, approved documents over chat snippets
3) Make it permission-aware end to end
This is critical for employee copilots.
Retrieval must respect:
- Document ACLs
- Group membership
- Org/region boundaries
- Role-based access
- Data sensitivity labels
Implementation pattern:
- Store ACL metadata with each chunk
- Apply permission filters before or during retrieval
- Never let the model “see” unauthorized text
- Log access for auditability
If permissions are complicated, use per-user filtered indexes or search-time ACL enforcement.
4) Chunk documents intelligently
Don’t chunk blindly by token count only.
Use chunking that preserves:
- Headings
- Section boundaries
- Tables/lists
- Parent-child relationships
Practical approach:
- Split by document structure first
- Keep chunks around ~300–800 tokens
- Add overlap only where needed
- Store titles, section headers, and breadcrumbs with each chunk
For policies and SOPs, section-aware chunks work much better than arbitrary slices.
5) Retrieve in stages
A strong pattern is:
-
Query understanding
- Detect intent, entities, doc type, and whether the user wants policy, how-to, status, or owner info
-
Candidate retrieval
- Hybrid search across indexed sources
-
Reranking
- Use a reranker or LLM-based scoring to sort the best passages
-
Context assembly
- Deduplicate, enforce permissions, include citations
-
Answer generation
- Require citations/links to source passages
This usually performs better than “top-k vector hits straight into the prompt.”
6) Ground answers in citations
Employee copilots need trust.
Return:
- Answer
- Source citations
- Document names
- Last updated dates
- Links to source docs
If the system can’t find support, it should say so explicitly:
- “I couldn’t find an authoritative policy on this.”
- “Here are the closest sources.”
What I’d choose for each retrieval type
For policies / process docs
- Hybrid search
- Section-aware chunking
- Strong reranking
- Canonical doc preference
- Freshness bias
For people / ownership / org info
- Structured lookup first
- Then search over directory/wiki docs
- Often better than pure vector search
For technical/internal knowledge
- Hybrid search over docs + tickets + runbooks
- Entity-aware indexing
- Strong metadata filtering by system/team
For Slack/Teams knowledge
- Use sparingly
- Index only approved channels
- Prefer threads with clear resolution
- Down-rank ephemeral content unless explicitly requested
A solid production architecture
Ingestion
- Connectors pull docs from source systems
- Normalize text, preserve structure
- Extract metadata: owner, source, timestamp, ACLs, doc type
Indexing
- Keyword index
- Vector index
- Optional graph/relationship index for owners, projects, systems
Retrieval
- Filter by permissions and metadata
- Run hybrid search
- Rerank results
- Fetch full relevant passages
Answering
- LLM gets only allowed, ranked evidence
- Produces answer with citations
- Fallback behavior if evidence is weak
Common mistakes to avoid
- Using only embeddings
- Ignoring permissions until after retrieval
- Chunking without document structure
- Mixing stale and canonical sources equally
- Letting the model answer without citations
- Including chat logs without provenance
- Not tuning for different query types
If you want a simple starting point
Start with:
- Hybrid search
- Permission filters
- Structured chunking
- Reranking
- Cited answers
- Canonical source preference
That’s the best balance of usefulness and safety for an employee copilot.
If you want, I can also propose:
- a concrete architecture diagram,
- an index/schema design, or
- a retrieval strategy by use case (HR, IT, legal, engineering, sales).