Prompt
How do I enforce data privacy rules on prompts before sending them to model providers?
Latest observation
To enforce data privacy rules before prompts leave your system, treat prompt handling like a policy enforcement pipeline rather than a simple string pass-through.
1) Define your privacy policy first
Start with clear rules such as:
- Denylist/Allowlist fields: what can never be sent
- e.g. SSNs, credit card numbers, API keys, health records
- Data classes: public, internal, confidential, regulated
- Destination rules:
- Which providers may receive which classes of data
- Which regions are allowed
- Purpose rules:
- Support chat may allow more than analytics
- Retention rules:
- Whether prompts can be logged, cached, or stored by the provider
A good policy is explicit, machine-readable, and versioned.
2) Put a policy engine in front of every outbound prompt
Before a prompt is sent to any model provider, run it through a pipeline like:
- Ingest
- Classify
- Detect sensitive data
- Transform/redact
- Approve/block
- Send to provider
Typical enforcement actions:
- Redact: replace sensitive values with placeholders
- Mask: partially hide values, e.g.
****1234 - Tokenize: replace with reversible internal tokens
- Block: stop the request entirely
- Route: send to a different provider/model that is approved for that data class
3) Detect sensitive data automatically
Use multiple detectors, not just regexes:
- Regex rules
- SSNs, card numbers, emails, phone numbers, secrets patterns
- DLP/PII classifiers
- Names, addresses, medical terms, financial identifiers
- Secret scanners
- API keys, bearer tokens, private keys, JWTs
- Custom entity recognizers
- Organization-specific IDs, customer numbers, project codenames
Best practice: combine:
- deterministic matching for high-confidence items
- ML/NLP classification for context-based detection
4) Redact or tokenize before transmission
Examples:
Redaction
Original:
“My SSN is 123-45-6789 and my email is jane@corp.com”
Sent:
“My SSN is [REDACTED_SSN] and my email is [REDACTED_EMAIL]”
Tokenization
Original:
“Customer 982173 has account 4411223344”
Sent:
“Customer [CUST_14] has account [ACCT_88]”
Keep the mapping inside your trust boundary if you need reversibility.
5) Minimize what you send
Only send the minimum necessary context:
- Strip metadata that isn’t needed
- Limit conversation history
- Summarize local context instead of sending full logs
- Remove attachments unless necessary
- Truncate long inputs with privacy checks before truncation
A common mistake is sending the entire thread when only the last user message is needed.
6) Separate trusted and untrusted context
Tag inputs by source:
- User input
- Internal system prompt
- Retrieved documents
- Tool outputs
- Database results
Then apply different rules to each source. For example:
- User input: sanitize for PII
- Tool output: may contain secrets; often stricter
- Retrieved docs: only approved excerpts allowed
7) Enforce provider-specific restrictions
Different providers have different privacy settings, so configure per vendor:
- Disable training on your data, if available
- Turn off or minimize logging
- Choose data residency/region settings
- Use enterprise or “no retention” modes when possible
- Restrict which endpoints can receive which data
Treat provider capabilities as part of policy enforcement, not just a procurement detail.
8) Add allowlisting for outbound fields
For structured prompts, prefer an allowlist over a denylist.
Example:
- Allowed:
customer_issue_summary,product_name,error_code - Not allowed:
full_name,email,phone,payment_info
If your application can generate a structured prompt object, validate it against a schema before rendering.
9) Log safely
You also need to protect the logs:
- Don’t log raw prompts by default
- Log hashes, IDs, or redacted versions
- Separate audit logs from application logs
- Restrict access to logs
- Set retention limits
A lot of privacy incidents happen in observability tooling, not the model provider.
10) Build a fail-closed system
If detection fails or the policy engine can’t classify the prompt:
- block the request
- require human review
- fall back to a safer model or local processing
Do not fail open for sensitive workflows.
A practical reference architecture
A robust setup looks like this:
App/UI → Prompt Builder → DLP/PII Scanner → Policy Engine → Redactor/Tokenizer → Provider Router → Model Provider
Optional:
- Audit logger
- Human approval queue
- Secrets vault for token maps
- Policy-as-code store
Example policy logic
You can express rules like:
- Block if any
PCI,SSN, orprivate_keyis present - Redact emails and phone numbers
- Allow product names and generic issue descriptions
- Only send HR data to approved internal model endpoint
- Do not send any data from EU users to non-EU endpoints
Implementation tips
- Use policy-as-code tools where possible:
- OPA/Rego, Cedar, custom policy rules
- Run checks at multiple layers:
- client-side, server-side, and pre-provider
- Version your policies and test them
- Create synthetic sensitive data test cases
- Measure false positives/negatives continuously
Example pseudocode
def prepare_prompt(prompt, context):
findings = scan_for_sensitive_data(prompt, context)
decision = evaluate_policy(findings, context)
if decision.action == "block":
raise PrivacyError(decision.reason)
sanitized = redact_or_tokenize(prompt, findings)
provider = route_provider(context, decision)
return sanitized, provider
Key principle
Never rely on the model provider to enforce your privacy rules for you.
Enforce them before the payload leaves your environment, and make the default behavior to minimize, redact, or block.
If you want, I can also provide:
- a sample policy schema,
- a Rego/OPA example, or
- a prompt sanitization pipeline in Python/TypeScript.