Prompt
How do I sync a RAG index from Google Drive and SharePoint?
Latest observation
To sync a RAG index from Google Drive and SharePoint, you generally want an incremental ingestion pipeline that:
- Connects to each source
- Detects changes (new/updated/deleted files)
- Re-chunks and re-embeds only changed content
- Upserts/deletes vectors in your RAG index on a schedule or webhook/event basis
Typical architecture
1) Connectors
Use a source-specific connector for each system:
- Google Drive
- Google Drive API
- Auth via OAuth or service account with domain-wide delegation
- SharePoint
- Microsoft Graph API
- Auth via Azure AD app registration / OAuth client credentials
Each connector should pull:
- file metadata
- content
- last modified time
- file ID / stable source ID
- parent path / folder metadata
- deletion state
2) Change detection
You need a sync strategy per source.
Google Drive
Use:
- Changes API to track incremental updates
- Or Drive push notifications / webhooks for near real-time updates
Track:
- file created/updated
- file deleted or trashed
- permission changes if access control matters
SharePoint
Use:
- Microsoft Graph delta queries for incremental sync
- Or webhooks/subscriptions for change notifications
Track:
- document library items added/updated/deleted
- folder path changes
- permission changes if you enforce ACL filtering
3) Normalize content
Before indexing, convert documents into a common internal format:
- plain text
- markdown
- OCR text if needed
- metadata fields
Typical metadata:
source:google_driveorsharepointsource_idtitleurlpathlast_modifiedmime_typeacl/ permissionsversion
4) Chunk and embed
For each changed document:
- extract text
- chunk it into passages
- embed each chunk
- store vectors with metadata linking them back to the document
Important:
- keep a stable
document_id - store
chunk_iddeterministically if possible - include source metadata for filtering and deletion
5) Upsert and delete
When a document changes:
- delete old chunks for that document
- insert/upsert new chunks
When a document is removed:
- delete all vectors associated with its
source_id
A common pattern is:
document_tablestores doc-level sync statevector_tablestores chunkssource_idmaps all chunks back to a file
Recommended sync flow
Initial full sync
- List all files in Drive/SharePoint
- Download and process each file
- Save a sync cursor / checkpoint
Ongoing incremental sync
- Poll changes API/delta API periodically
- e.g. every 5–15 minutes
- For each changed item:
- if updated: reprocess and reindex
- if deleted: remove vectors
- Advance cursor/checkpoint after success
Best practices
Use source cursors
Persist:
- Google Drive
startPageToken/ changes token - SharePoint Graph
deltaLink
So you can resume sync without re-scanning everything.
Handle ACLs
If users should only retrieve documents they can access:
- store permissions metadata
- filter retrieval by user/group ACLs at query time
Deduplicate
Files may be duplicated across shared locations or re-uploaded:
- hash normalized content
- use source IDs plus content hash to avoid redundant embeddings
Versioning
Keep version history if useful:
- current version only
- or historical versions with
version_id
Large files
For PDFs, Office docs, or scanned images:
- extract text with document parsers
- OCR images if needed
Example implementation pattern
Sync worker pseudo-flow
for each source in [google_drive, sharepoint]:
cursor = load_cursor(source)
changes = fetch_changes(source, cursor)
for change in changes:
if change.deleted:
delete_vectors(change.source_id)
mark_deleted(change.source_id)
else:
doc = download_and_parse(change)
chunks = chunk(doc.text)
vectors = embed(chunks)
upsert_vectors(change.source_id, vectors, metadata)
save_cursor(source, changes.new_cursor)
Tools that can help
Depending on your stack:
- LlamaIndex / LangChain for document loading and chunking
- Unstructured for parsing office docs / PDFs
- Pinecone, Weaviate, Qdrant, Milvus, pgvector for vector storage
- Airbyte / Fivetran / custom ETL for ingestion pipelines
- Microsoft Graph SDK and Google Drive SDK for connectors
If you want near-real-time sync
Use:
- Drive push notifications + changes API
- SharePoint Graph subscriptions + delta queries
A common pattern is:
- webhook triggers a job
- job calls delta/changes API to fetch exact changes
- indexer processes updates
This avoids missing events and makes sync reliable.
Practical recommendation
If you’re building this yourself, the simplest reliable setup is:
- Google Drive: Changes API + scheduled polling
- SharePoint: Microsoft Graph delta queries + scheduled polling
- Vector DB: upsert/delete by
source_id - Metadata store: track cursor + doc hashes + ACLs
If you want, I can also give you:
- a reference architecture diagram,
- sample Python code for Drive + SharePoint sync, or
- a LangChain/LlamaIndex ingestion example.