Prompt
I'm building a multi-tenant AI app and need auth and isolation for model serving
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
For a multi-tenant AI app, the key is to treat authn/authz and data/model isolation as separate but connected layers.
Core goals
- Authenticate users/services
- Authorize per-tenant access
- Isolate tenant data, prompts, embeddings, logs, and model outputs
- Prevent cross-tenant leakage in serving infrastructure
- Audit every sensitive action
Recommended architecture
1) Identity and authentication
Use a centralized IdP:
- OIDC/OAuth2 for users
- mTLS or signed service tokens for internal services
- Short-lived access tokens
- Rotate signing keys regularly
Common setup:
- User logs in via SSO
- IdP issues JWT with:
sub= user IDtenant_idrolesscopes- token expiry
2) Tenant-aware authorization
Every request must carry a tenant context and be checked against it.
Rules:
- Never trust tenant ID from the client alone
- Derive tenant from validated JWT/session
- Enforce row-level / object-level authorization
- Use policy checks at API gateway and again in backend services
Typical approach:
tenant_idclaim in token- Backend middleware injects
tenant_idinto request context - All DB queries filtered by
tenant_id - All storage paths namespaced by tenant
3) Data isolation patterns
Choose one based on risk and scale:
Option A: Shared DB, tenant-scoped rows
- Every table includes
tenant_id - Enforce with:
- application checks
- DB row-level security if supported
- Good for many tenants, lower cost
Option B: Separate schema per tenant
- Moderate isolation
- Better than shared rows
- More operational overhead
Option C: Separate DB per tenant
- Stronger isolation
- Easier compliance boundaries
- Higher cost/ops
For AI apps, a common hybrid:
- Shared control plane
- Separate tenant data stores for sensitive customers
- Shared inference layer with strict request isolation
4) Model-serving isolation
This is where AI apps often leak.
Important boundaries:
- Per-request context isolation
- Per-tenant memory isolation
- No shared prompt/state across tenants
- No shared caches without tenant keying
- No cross-tenant retrieval from vector DB
For inference:
- Use stateless model servers if possible
- If stateful, reset state between requests
- Namespace:
- prompt cache
- conversation history
- embeddings
- feature store
- tool/function execution context
Vector DB / RAG:
- Partition by tenant
- Add
tenant_idmetadata to every chunk - Filter retrieval by tenant before similarity search if possible
- Verify that top-k results are only from the same tenant
5) Network and runtime isolation
- Separate namespaces or clusters for high-sensitivity tenants
- Kubernetes:
- namespace per tenant or per tier
- network policies
- service accounts with least privilege
- Use secrets managers, not env vars in plain text
- Encrypt in transit and at rest
6) Prompt and output safety
- Treat prompts as sensitive data
- Log carefully; redact secrets/PII
- Avoid storing raw prompts unless needed
- If using tool calls, validate all tool inputs and outputs
- Prevent model from accessing tenant-unrelated data sources
7) Audit and observability
Log:
- authenticated principal
- tenant ID
- resource accessed
- model/version used
- retrieval sources
- tool calls
- timestamps
- allow/deny decisions
Use:
- tamper-evident logs
- per-tenant audit views
- anomaly detection for unusual access patterns
Practical implementation pattern
Request flow
- User signs in with IdP
- App receives JWT with
tenant_id - API gateway validates token
- Backend checks RBAC/ABAC policy
- Request goes to inference service with tenant context
- Retrieval layer filters tenant data
- Model server runs stateless inference
- Output is returned and logged with tenant-scoped audit
Example policy idea
Allow access only if:
- token valid
token.tenant_id == resource.tenant_id- user has required role
- resource not suspended
Security checklist
- JWT validation with exp/aud/iss checks
- Tenant ID enforced server-side
- Row-level security or equivalent
- Vector search tenant filtering
- Separate encryption keys per tenant or tier
- Least-privilege service accounts
- No shared mutable model state
- Audit logs with tenant context
- Secrets manager and key rotation
- Rate limiting per tenant
- Abuse detection and revocation
Recommended default design
If you want a balanced starting point:
- OIDC login
- JWT with tenant_id
- Shared app services
- Shared DB with row-level security
- Tenant-filtered vector store
- Stateless model serving
- Per-tenant encryption keys for sensitive data
- Strict audit logs
If you want, I can also give you:
- a reference architecture diagram
- a JWT claims schema
- a Kubernetes isolation pattern
- a sample policy model for RBAC/ABAC