Prompt
I'm building a copilot over emails, docs, and database tables. How should I connect all the sources?
Latest observation
A good way to think about this is: don’t “connect” everything directly to the model. Instead, build a unified retrieval + action layer that sits between the LLM and your sources.
Recommended architecture
1) Ingest each source into a common index
For each system:
- Emails: pull message text, subject, sender, recipients, timestamps, threads, attachments
- Docs: pull document text, headings, metadata, links, version history
- Database tables: pull schema, table/column descriptions, sample rows, and either:
- sync full rows for smaller tables, or
- expose live query tools for larger/fast-changing tables
Normalize everything into a shared record format like:
{
"source_type": "email | doc | table_row | table_schema",
"source_id": "...",
"title": "...",
"text": "...",
"metadata": {
"author": "...",
"created_at": "...",
"updated_at": "...",
"permissions": ["..."]
}
}
2) Use two retrieval paths
You usually want both:
- Semantic search over embeddings for fuzzy lookup
- Structured filtering over metadata for exact constraints
For example:
- “What did Sarah say last week about pricing?” → semantic search in emails
- “Show all invoices over $10k in Q3” → database query/tool call
3) Keep source-specific connectors, but unify the interface
Build adapters for each source:
EmailConnectorDocsConnectorDatabaseConnector
Each connector should expose a common contract like:
search(query, filters)fetch(id)list_changes(since_cursor)get_permissions(user)
Then your copilot can work against one abstraction, while the connector handles Gmail/Outlook, SharePoint/Drive/Notion/Confluence, Postgres/Snowflake/etc.
4) Put permissions first
This is critical. Retrieval must be filtered by the end user’s access:
- propagate ACLs from each source
- enforce row/document/email-level security at retrieval time
- never let the LLM see unauthorized content
A common pattern:
- index only content the user is allowed to access
- or index everything with ACL metadata and filter before returning results
5) Separate “search” from “action”
For databases especially:
- use retrieval for understanding context
- use tools for live operations:
- SQL query execution
- report generation
- record updates
- ticket creation
Don’t rely on embeddings alone for tables; use the DB as a system of record.
Practical pattern by source
Emails
Best handled as:
- thread-aware chunking
- metadata-rich indexing
- embedding over message body + subject
- keep sender/recipient/time/thread IDs
Docs
Best handled as:
- chunk by section/heading
- preserve hierarchy
- store links and references
- re-index on edits/version changes
Tables
Use a hybrid:
- index schema and business definitions for semantic understanding
- query the live DB for exact numbers and up-to-date results
- optionally cache materialized extracts for common analytics
How the copilot should work at runtime
- User asks a question
- LLM classifies intent:
- search emails
- search docs
- query tables
- or combine them
- Retrieval layer fetches relevant context from each source
- LLM synthesizes answer with citations
- If needed, LLM calls a tool to take action
Best practice: don’t flatten everything too much
You want a common retrieval layer, but still preserve source-specific structure:
- emails have threads and conversational context
- docs have hierarchy and versioning
- tables have schema, joins, and live data
If you flatten them all into plain text too early, you lose valuable structure.
A strong starting stack
A common setup is:
- Connectors: Gmail/Outlook, Google Drive/Confluence/Notion, Postgres/Snowflake/BigQuery
- Index: vector DB + keyword index + metadata store
- Orchestration: tool-calling agent
- Permissions: ACL filter service
- Answering: RAG with citations and source links
If you want the simplest version first
Start with:
- one vector index for docs/emails
- live SQL tool for tables
- metadata filters for permissions
- citations back to source records
Then add:
- incremental sync
- thread/section-aware chunking
- hybrid keyword + vector search
- cross-source joins at the retrieval layer
If you want, I can sketch a reference architecture diagram or give you a data model and API interface for the connectors.