Prompt
I’m building a search product for sensitive company content. How do I avoid permission leaks?
Latest observation
To avoid permission leaks in a search product for sensitive company content, design security trimming as a first-class feature, not a last-mile filter.
Core principles
-
Filter before ranking, not after
- Never rank across all documents and then hide unauthorized results.
- Build the candidate set only from items the user is allowed to see.
-
Evaluate permissions at query time
- Determine authorization using the current user identity, groups, roles, tenant, and any document-level ACLs.
- Don’t trust cached results unless they’re keyed to the exact authorization context.
-
Use the same auth source as the source system
- Ingest ACLs from the authoritative system of record.
- Keep them synced continuously or near-real-time.
- If permissions change, revoke access in search immediately or very quickly.
-
Treat permissions as data, not metadata
- Store ACLs in a way the search engine can efficiently filter on.
- Common pattern: index allow/deny principals and tenant IDs on each document.
-
Default deny
- If permissions are missing, malformed, or stale, do not show the document.
Common implementation patterns
A. Security-trimmed indexing
At ingest time, attach fields like:
tenant_idallowed_usersallowed_groupsdenied_usersdenied_groupsclassification_level
At query time:
- build an authorization predicate for the user
- apply it as a mandatory filter
This is the most common approach for enterprise search.
B. Separate indexes by tenant or sensitivity tier
Useful if your tenancy model is strict.
- One index per tenant, or
- one index per sensitivity class
This reduces blast radius, but you still need document-level ACLs in many cases.
C. Precomputed access tokens
Instead of joining against a permission database for every query:
- compute a user’s effective access tokens
- index those tokens on documents
- filter by token overlap
This can be fast, but watch out for:
- token explosion
- stale group membership
- privilege inheritance complexity
Things that commonly cause permission leaks
-
Post-filtering results after ranking
- Unauthorized docs can influence ranking or snippets.
-
Shared caches without auth keys
- Query results cached for one user are served to another.
-
Snippet/highlight leakage
- Even if the title is hidden, a snippet can reveal sensitive text.
-
Autocomplete and suggestions
- Suggestions can leak document names, terms, or project codenames.
-
Aggregations/facets
- Counts can reveal the existence of restricted content.
- Facet values can leak sensitive categories.
-
Recently viewed / trending / related items
- These features often ignore ACLs unless explicitly trimmed.
-
Search logs and analytics
- Logs may store raw queries, titles, snippets, or doc IDs. Protect them too.
-
Deleted or revoked access not invalidated
- Search results stay visible after source permissions change.
Hardening checklist
- Enforce ACLs in the search backend, not just the UI
- Apply ACL filters before scoring/ranking
- Include tenant isolation in every auth check
- Deny by default if ACL data is absent or stale
- Sync permission changes quickly
- Security-trim:
- results
- snippets
- autosuggest
- facets
- related items
- analytics exports
- Key caches by user/session/authorization fingerprint
- Test with adversarial cases:
- removed group membership
- nested groups
- shared documents
- revoked access mid-session
- cross-tenant queries
Good test cases to run
-
User A and User B search the same term.
- Verify A never sees B’s restricted docs or snippets.
-
Remove a user from a group.
- Verify access disappears immediately or within your SLA.
-
Search for a term that exists only in a restricted document.
- Verify no result counts, suggestions, or snippets reveal it.
-
Same query from different tenants.
- Verify zero cross-tenant leakage in results, facets, or caches.
Practical recommendation
If you’re early in the build, the safest architecture is:
- ingest documents with ACLs from the source system
- compute effective user principals at query time
- filter in the search engine using those principals
- security-trim all derived features
- make revocation fast and observable
If you want, I can also give you:
- a reference architecture,
- a permission model schema,
- or an example query flow for Elasticsearch/OpenSearch/SQL-based search.