Prompt
I need a compliant way to detect PII leakage in LLM outputs
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
Yes — the safest compliant pattern is to treat PII detection as a separate post-processing control, not something you rely on the model to do by itself.
A compliant architecture
1) Define what counts as PII
Start with a policy-backed taxonomy, for example:
- Direct identifiers: name, email, phone, SSN/national ID, passport, account numbers
- Quasi-identifiers: date of birth, ZIP/postcode, IP address, device ID
- Sensitive data: health, financial, biometrics, religion, etc.
- Organization-specific secrets: API keys, access tokens, customer IDs
Map these to the laws/regulations that matter for your use case:
- GDPR / UK GDPR
- CCPA/CPRA
- HIPAA
- PCI DSS
- Local data protection laws
2) Detect PII in the model output before release
Use a layered detector, not a single regex:
Layer A: deterministic rules
- Regex for emails, phone numbers, SSNs, credit cards, IBANs
- Luhn check for payment cards
- Pattern checks for API keys / tokens / secrets
- Dictionaries for common identifiers where applicable
Layer B: NLP / ML entity detection
- Named Entity Recognition for PERSON, ORG, LOC, DATE, etc.
- Domain-specific models for medical, legal, financial identifiers
Layer C: policy logic
- Flag if multiple quasi-identifiers appear together
- Flag if output resembles a user record, account lookup, or profile
- Flag if the output contains low-frequency unique strings copied from input
Layer D: contextual risk scoring
- Distinguish between:
- user-supplied PII repeated back
- model-generated hallucinated PII
- system prompt leakage
- memorized training data leakage
3) Redact or block based on policy
Have an action policy, for example:
- Block: direct identifiers, secrets, payment card data
- Redact: names, emails, phone numbers, addresses
- Allow with warning: generalized or already-public info
- Escalate: if the output contains multiple sensitive elements or high-confidence secrets
4) Keep the check outside the LLM
For compliance, don’t ask the model “does this contain PII?” as the only control. Instead:
- Generate output
- Scan output with independent detectors
- Apply redaction/blocking
- Log the decision and detector version
5) Add input/output data minimization
To reduce leakage risk:
- Don’t send unnecessary PII to the model
- Mask identifiers before prompting where possible
- Use retrieval filters so the model never sees irrelevant sensitive fields
- Restrict long context windows containing raw personal data
6) Maintain auditability
For compliance, you usually want:
- Versioned PII policy
- Detector version and threshold records
- Alert logs
- Human review workflow
- Retention limits on scanned text and logs
Practical detection stack
A common stack is:
- Regex + checksum validators
- NER / entity extraction
- Secret scanners for keys/tokens
- Custom rules for your domain
- Human review for edge cases
Examples of tooling categories:
- PII/entity detection libraries
- Secret scanning tools
- DLP engines
- Cloud content moderation / DLP services
Important compliance notes
- If you process personal data to detect PII, that processing itself may be subject to privacy rules, so document purpose and retention.
- Minimize storage of raw outputs; store hashes or redacted versions where possible.
- Make sure vendor contracts and data processing agreements cover any external scanning services.
- Calibrate false positives and false negatives; compliance usually favors conservative blocking for sensitive categories.
Recommended implementation pattern
A simple policy pipeline:
- Generate response
- Run deterministic detectors
- Run entity/PII classifier
- Combine scores
- If sensitive:
- redact or block
- log reason
- optionally send to human review
- Return sanitized output
Example policy
- Email/phone/SSN/credit card: always redact
- Names and addresses: redact unless the user explicitly requested them and policy allows
- Secrets/tokens/credentials: always block and alert
- Health or financial data: always escalate to review
If you want, I can help you with either:
- a compliance-oriented detection policy template, or
- a reference implementation in Python for post-processing LLM outputs.