Prompt
I need a way to detect PII leakage in model outputs
Latest observation
A practical way to detect PII leakage in model outputs is to use a layered detector rather than a single regex. The best setup usually combines:
-
Pattern checks for obvious PII
- Email addresses
- Phone numbers
- SSNs / national IDs
- Credit cards
- IP addresses
- Postal addresses
- Dates of birth
- Account numbers
-
Named entity recognition (NER)
Detect people, organizations, locations, and especially combinations like:- person name + phone number
- person name + address
- employee name + internal identifier
-
Contextual rules
Catch cases where something is not always PII by itself but is sensitive in context:- “John Smith, SSN …”
- “Here is the customer’s private token…”
- “My medical record number is …”
-
Fuzzy matching against known sensitive data
If you have a protected dataset, compare outputs against:- user names
- customer records
- internal IDs
- conversation memory
- retrieved documents
-
LLM-based review for ambiguous cases
Use a second model to classify whether text contains PII, especially for free-form responses where patterns miss things like:- “You can reach my sister at her work number…”
- “The address is the blue house across from…”
A solid implementation approach
Option A: Rule-based + NER + scoring
Assign a leakage score to each output:
- +3 for email/phone/SSN/credit card
- +2 for full name, address, DOB
- +1 for account-like identifiers or location hints
- +2 if multiple PII types appear together
Trigger an alert if score exceeds a threshold.
Option B: Use a PII detection library
Some common tools:
- Microsoft Presidio — strong open-source PII detection framework
- spaCy — for NER, often combined with custom rules
- Hugging Face token classification models — for custom PII models
Presidio is a good starting point because it supports:
- regex detectors
- NLP-based recognizers
- custom recognizers
- redaction and anonymization
Example: Python with Presidio
from presidio_analyzer import AnalyzerEngine
analyzer = AnalyzerEngine()
text = "Contact Jane Doe at jane.doe@example.com or +1 (415) 555-1234."
results = analyzer.analyze(text=text, language="en")
for r in results:
print(r.entity_type, r.start, r.end, text[r.start:r.end], r.score)
Example policy
You can classify outputs as:
- Safe: no PII detected
- Needs review: possible PII or low-confidence match
- Leakage: confirmed PII present
Best practices
- Normalize text before scanning:
- decode HTML entities
- remove zero-width characters
- handle spaced-out emails/phones like
j o h n @ ...
- Scan both:
- raw model output
- tool outputs / retrieved context
- Use allowlists for benign patterns:
- product names that look like names
- example email addresses like
example.com
- Log detections with:
- detected type
- confidence
- source prompt/output
- action taken
If you want production-grade protection
Combine:
- input filtering: prevent the model from seeing unnecessary sensitive data
- output filtering: scan every response before delivery
- retrieval filtering: scan RAG context
- policy enforcement: block or redact when PII is detected
If you want, I can give you:
- a ready-to-use Python detector,
- a regex-based quick solution, or
- a PII detection policy for LLM outputs.