Prompt
How do I build a retrieval layer that handles fresh documents and permissions?
Latest observation
A good retrieval layer for fresh documents + permissions usually needs to solve 4 problems at once:
- Ingest new/updated/deleted docs quickly
- Index them in a search-friendly way
- Filter results by user permissions at query time
- Keep permissions and index state consistent enough to trust
Here’s a practical architecture that works well.
1) Use an ingestion pipeline, not direct indexing
Don’t have the app write directly to the search index. Instead:
- Source of truth: database / object store / CMS / file system
- Change event: document created/updated/deleted
- Queue / stream: Kafka, SQS, Pub/Sub, Redis streams, etc.
- Indexer workers:
- fetch the latest doc
- extract text/metadata
- chunk it
- enrich it
- write to the retrieval index
This gives you:
- retries
- backpressure control
- better observability
- easier reindexing
2) Store permissions as first-class metadata
For each document, store access control data in the index alongside the content.
Common patterns:
A. ACL lists
Store:
allowed_usersallowed_groupstenant_idvisibility
Then at query time, filter on the requesting user’s groups and user ID.
Example fields:
{
"doc_id": "123",
"tenant_id": "acme",
"title": "Q4 roadmap",
"content": "...",
"allowed_users": ["u1", "u2"],
"allowed_groups": ["g9", "g12"],
"public": false
}
B. Security labels / roles
Instead of explicit ACLs, store labels like:
confidentialhr_onlyengineeringproject_x
Then map user entitlements to labels.
C. Hybrid
Common in enterprise search:
- tenant isolation
- document ACLs
- group-based permissions
- classification labels
3) Filter at retrieval time, not after
A common mistake is:
- retrieve top 50 docs
- filter unauthorized docs
- return whatever’s left
This can hurt relevance and leak signals. Better:
- push the permission filter into the search query itself
- only rank authorized documents
Most search engines support this via metadata filters.
Example conceptually:
WHERE tenant_id = :tenant
AND (
public = true
OR :user_id IN allowed_users
OR allowed_groups INTERSECT :user_groups IS NOT EMPTY
)
In vector search, use:
- pre-filtering if supported
- hybrid search with filters
- or overfetch + secure re-rank if the engine can’t filter natively
4) Separate document freshness from query freshness
Freshness means different things:
Content freshness
New/updated doc appears in search quickly.
Permission freshness
If access changes, that change must reflect quickly too.
Deletion freshness
Deleted docs should disappear promptly.
To handle this:
- emit events for
document.updated,document.deleted,permissions.changed - reindex documents on both content and ACL changes
- include a
versionorupdated_atto ignore stale events
A robust pattern:
- store a
doc_version - indexers only apply an event if it’s the latest version
- on delete, tombstone the doc in the index immediately
5) Build a permission cache or entitlement service
For large orgs, computing group memberships or ACL expansion on every query can be expensive.
Use a service or cache that answers:
- user’s groups
- roles
- tenant
- derived entitlements
Then the retrieval layer gets a compact permission token, such as:
{
"user_id": "u1",
"tenant_id": "acme",
"groups": ["g9", "g12"],
"roles": ["editor"]
}
If permissions are complex, precompute:
- transitive group membership
- project memberships
- resource-level entitlements
- per-user ACL bitsets or tokens
6) Index documents in chunks, but keep doc-level ACLs
For semantic retrieval, you’ll often chunk documents.
Best practice:
- each chunk is a searchable unit
- each chunk inherits the parent document’s permissions
- store:
doc_idchunk_idchunk_textacl fields
That way you can retrieve relevant chunks while preserving doc security.
7) Support incremental updates and reindexing
You need two flows:
Incremental path
For normal updates:
- reindex only changed docs/chunks
Backfill path
For schema changes, bug fixes, model changes:
- full reindex from source of truth
Make reindex idempotent:
- same input event shouldn’t create duplicates
- use deterministic IDs like
doc_id + chunk_id
8) Use soft deletes or tombstones
When docs are deleted:
- mark them deleted in source of truth
- emit delete event
- remove from index or mark inactive
- prevent retrieval from returning stale docs
Soft deletes help with eventual consistency:
is_deleted = true- filter them out immediately
- later compact them away
9) Design for eventual consistency and race conditions
There will be short windows where:
- document content is updated but index isn’t
- permissions changed but old ACL still exists in index
- delete event arrives before update event
To manage this:
- include
version/updated_at - ignore out-of-order events
- periodically reconcile index vs source of truth
- run consistency checks
For high-security systems, consider:
- permission checks against an authoritative service at query time, in addition to indexed metadata
10) Recommended query flow
A strong pattern is:
- Authenticate user
- Get user entitlements
- Build query with permission filters
- Search index
- Re-rank results if needed
- Return only authorized results
Example:
- query text: “budget for project alpha”
- filters:
- tenant = acme
- allowed by ACL / group membership
- retrieve top 20
- semantic rerank
- return top 5
11) If you use vector search, watch out for permission leakage
Vector DBs often support metadata filters, but you need to ensure:
- filters are applied before final selection, not after
- chunk metadata includes ACLs
- embeddings are regenerated when content changes materially
- sensitive docs don’t leak through summaries or nearby chunks
If the vector store cannot filter securely:
- use a two-stage system:
- candidate retrieval from vector DB
- secure authorization filter
- rerank only authorized candidates
12) Practical minimal architecture
A good baseline stack:
- Source of truth: Postgres / SharePoint / S3 / CMS
- Event stream: Kafka / SQS
- Indexer service: parses + chunks + embeds
- Search store: OpenSearch / Elasticsearch / pgvector / Pinecone / Weaviate
- Permission service: resolves user groups/roles/ACLs
- Retrieval API: authenticates, filters, queries, reranks
13) Things to avoid
- indexing without ACL metadata
- filtering permissions only in the app after retrieval
- using stale group membership data with no refresh plan
- relying on full reindexing for every update
- mixing tenants in one index without strong isolation
- returning raw chunks that can reveal unauthorized context
14) A simple rule of thumb
If you remember only one thing:
Treat permissions as searchable metadata and freshness as an event-driven indexing problem.
That combination solves most retrieval-layer issues.
If you want, I can also give you:
- a reference architecture diagram
- a Postgres + OpenSearch implementation
- a vector DB design for permission-aware RAG
- or a sample schema and query examples