Prompt

I'm building an internal copilot over SharePoint and Google Drive. What retrieval setup should I use?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

For an internal copilot over SharePoint + Google Drive, I’d recommend a hybrid, source-aware retrieval architecture rather than “one vector index for everything.”

Recommended setup

1) Ingest into a unified document pipeline

Pull content from both sources into a common schema:

  • doc_id
  • source = SharePoint / Google Drive
  • source_url
  • title
  • owner
  • created_at, updated_at
  • permissions / ACLs
  • folder/path
  • content
  • content_type (doc, ppt, pdf, sheet, etc.)
  • version

This makes downstream retrieval much easier.


2) Use hybrid retrieval

Use both:

  • BM25 / keyword search for exact terms, project names, acronyms, policy IDs
  • Vector search for semantic similarity

For internal enterprise content, hybrid search is usually much better than pure embeddings alone.


3) Chunk documents intelligently

Split by structure, not just fixed length:

  • Headings
  • Paragraphs
  • Bullet lists
  • Table sections

Good defaults:

  • chunk size: ~300–800 tokens
  • overlap: ~50–150 tokens
  • preserve document hierarchy and section titles in each chunk

For PowerPoints and docs, include slide/page context in the chunk text.
For spreadsheets, index row groups or sheet summaries rather than raw cells only.


4) Store metadata separately and filter hard

Before semantic retrieval, enforce permissions filtering:

  • user can only retrieve documents they have access to
  • filter by source, team, department, date, doc type if needed

This is critical for SharePoint/Drive because ACLs vary by file/folder.


5) Use a two-stage retrieval flow

A strong pattern is:

  1. Candidate retrieval

    • hybrid search over chunks
    • permission-filtered
    • fetch top 50–200 candidates
  2. Reranking

    • use a cross-encoder / reranker
    • rerank to top 5–15 chunks

This improves relevance a lot, especially for enterprise docs with similar wording.


6) Consider document-level + chunk-level indexes

Use both:

  • Document-level index for finding the right file
  • Chunk-level index for finding the right passage

This helps when:

  • user asks “What is the latest vacation policy?”
  • you need to identify the policy doc first, then extract the relevant section

7) Add freshness-aware ranking

For internal content, recency often matters.

Boost:

  • newer versions
  • recently modified docs
  • canonical docs (policies, handbooks, source-of-truth pages)

But don’t let recency override authority if a policy doc is older but canonical.


8) Include source attribution in answers

Always return:

  • doc title
  • source platform
  • link
  • quoted passage or referenced section

This builds trust and makes it easy for users to verify.


Best-practice architecture

If I were designing this today, I’d use:

  • Unified ingestion + normalization
  • ACL-aware metadata filtering
  • Hybrid retrieval (BM25 + embeddings)
  • Cross-encoder reranking
  • Doc + chunk indexes
  • Freshness / authority boosting
  • Citations in final output

Practical stack options

If you want a managed approach

  • Elastic / OpenSearch hybrid search
  • Azure AI Search if you’re already in Microsoft ecosystem
  • Vertex AI Search if you’re more on Google Cloud

If you want to build it yourself

  • Ingestion: custom connectors to SharePoint + Google Drive
  • Index: Elasticsearch/OpenSearch + vector index, or a vector DB plus keyword engine
  • Reranker: sentence-transformers cross-encoder or hosted rerank API
  • Orchestration: LlamaIndex / LangChain / custom retrieval service

Important SharePoint/Drive specifics

SharePoint

  • Watch for nested permissions inherited from sites, libraries, folders, files
  • Preserve version history where useful
  • Handle Office docs with embedded tables and comments carefully

Google Drive

  • Use file ACLs and shared drive memberships
  • Be careful with shortcuts, duplicates, and “latest” file versions
  • Extract text from Docs/Slides/Sheets properly, not just exported PDFs

My default recommendation

If you want the most reliable setup:

Permission-filtered hybrid search + chunk-level indexing + reranking + citations

That’s the sweet spot for enterprise internal copilots.

If you want, I can also propose:

  1. a reference architecture diagram,
  2. a chunking/indexing strategy by file type, or
  3. a minimal production-ready stack for your cloud environment.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.