Prompt
I'm building an internal copilot over SharePoint and Google Drive. What retrieval setup should I use?
Latest observation
For an internal copilot over SharePoint + Google Drive, I’d recommend a hybrid, source-aware retrieval architecture rather than “one vector index for everything.”
Recommended setup
1) Ingest into a unified document pipeline
Pull content from both sources into a common schema:
doc_idsource= SharePoint / Google Drivesource_urltitleownercreated_at,updated_atpermissions/ ACLsfolder/pathcontentcontent_type(doc, ppt, pdf, sheet, etc.)version
This makes downstream retrieval much easier.
2) Use hybrid retrieval
Use both:
- BM25 / keyword search for exact terms, project names, acronyms, policy IDs
- Vector search for semantic similarity
For internal enterprise content, hybrid search is usually much better than pure embeddings alone.
3) Chunk documents intelligently
Split by structure, not just fixed length:
- Headings
- Paragraphs
- Bullet lists
- Table sections
Good defaults:
- chunk size: ~300–800 tokens
- overlap: ~50–150 tokens
- preserve document hierarchy and section titles in each chunk
For PowerPoints and docs, include slide/page context in the chunk text.
For spreadsheets, index row groups or sheet summaries rather than raw cells only.
4) Store metadata separately and filter hard
Before semantic retrieval, enforce permissions filtering:
- user can only retrieve documents they have access to
- filter by source, team, department, date, doc type if needed
This is critical for SharePoint/Drive because ACLs vary by file/folder.
5) Use a two-stage retrieval flow
A strong pattern is:
-
Candidate retrieval
- hybrid search over chunks
- permission-filtered
- fetch top 50–200 candidates
-
Reranking
- use a cross-encoder / reranker
- rerank to top 5–15 chunks
This improves relevance a lot, especially for enterprise docs with similar wording.
6) Consider document-level + chunk-level indexes
Use both:
- Document-level index for finding the right file
- Chunk-level index for finding the right passage
This helps when:
- user asks “What is the latest vacation policy?”
- you need to identify the policy doc first, then extract the relevant section
7) Add freshness-aware ranking
For internal content, recency often matters.
Boost:
- newer versions
- recently modified docs
- canonical docs (policies, handbooks, source-of-truth pages)
But don’t let recency override authority if a policy doc is older but canonical.
8) Include source attribution in answers
Always return:
- doc title
- source platform
- link
- quoted passage or referenced section
This builds trust and makes it easy for users to verify.
Best-practice architecture
If I were designing this today, I’d use:
- Unified ingestion + normalization
- ACL-aware metadata filtering
- Hybrid retrieval (BM25 + embeddings)
- Cross-encoder reranking
- Doc + chunk indexes
- Freshness / authority boosting
- Citations in final output
Practical stack options
If you want a managed approach
- Elastic / OpenSearch hybrid search
- Azure AI Search if you’re already in Microsoft ecosystem
- Vertex AI Search if you’re more on Google Cloud
If you want to build it yourself
- Ingestion: custom connectors to SharePoint + Google Drive
- Index: Elasticsearch/OpenSearch + vector index, or a vector DB plus keyword engine
- Reranker: sentence-transformers cross-encoder or hosted rerank API
- Orchestration: LlamaIndex / LangChain / custom retrieval service
Important SharePoint/Drive specifics
SharePoint
- Watch for nested permissions inherited from sites, libraries, folders, files
- Preserve version history where useful
- Handle Office docs with embedded tables and comments carefully
Google Drive
- Use file ACLs and shared drive memberships
- Be careful with shortcuts, duplicates, and “latest” file versions
- Extract text from Docs/Slides/Sheets properly, not just exported PDFs
My default recommendation
If you want the most reliable setup:
Permission-filtered hybrid search + chunk-level indexing + reranking + citations
That’s the sweet spot for enterprise internal copilots.
If you want, I can also propose:
- a reference architecture diagram,
- a chunking/indexing strategy by file type, or
- a minimal production-ready stack for your cloud environment.