Prompt
How do I add safety filters to prompts and outputs in an AI gateway?
Latest observation
To add safety filters to both prompts and outputs in an AI gateway, use a layered approach: classify → enforce → log → review. A gateway is a good place to centralize this because it sits between your apps and model providers.
1) Define what you want to block or transform
Start with a clear policy for both inbound and outbound traffic:
-
Prompts (input)
- PII: SSNs, credit cards, API keys, passwords
- Prompt injection: “ignore previous instructions,” “reveal system prompt”
- Harmful requests: malware, fraud, self-harm instructions, hate/harassment
- Data exfiltration: requests to expose secrets, internal docs
-
Outputs
- PII leakage
- Unsafe advice or disallowed content
- Policy violations: copyrighted text, confidential data, unsafe code
- Hallucinated sensitive claims if your domain requires accuracy controls
2) Put filters in two places
Use both pre-processing and post-processing:
A. Input filtering before the model
Do this as close to the gateway edge as possible:
- Pattern matching / regex for obvious secrets and PII
- DLP/PII detection service
- Prompt injection classifier
- Topic moderation classifier
- Context-aware allow/deny rules by user role, tenant, or endpoint
Actions:
- Block high-risk requests
- Redact sensitive parts
- Warn / require confirmation
- Route to a safer model or stricter policy tier
B. Output filtering after the model
Before returning the response:
- Scan for PII/secrets
- Run moderation/safety classification
- Check for policy-specific forbidden content
- Detect if the model echoed hidden instructions or internal data
Actions:
- Replace risky spans with redaction
- Refuse and return a safe completion
- Truncate if necessary
- Escalate to human review for borderline cases
3) Use multiple detection methods
No single filter is enough. Combine:
Deterministic checks
Good for:
- API keys
- JWTs
- SSNs
- Credit cards
- Known secret patterns
ML-based classifiers
Good for:
- Toxicity
- Hate/harassment
- Sexual content
- Self-harm
- Prompt injection
- Sensitive intent
Policy/rule engine
Good for:
- Tenant-specific rules
- Role-based access control
- Endpoint-specific restrictions
- Region/compliance constraints
4) Protect against prompt injection
Treat user content as untrusted. For gateway design:
- Separate system instructions from user input
- Never let user content overwrite policy
- Strip or flag meta-instructions like:
- “ignore previous instructions”
- “reveal your prompt”
- “act as system”
- If your app uses RAG, filter retrieved documents too, because they can carry injection text
Helpful practices:
- Mark retrieved content as untrusted
- Summarize or extract facts before sending to the model
- Use tool/function calling with strict schemas
- Limit what tools the model can access
5) Redact instead of raw block when possible
For usability, don’t always deny the request. Options:
- Replace detected secrets with
[REDACTED] - Mask PII partially
- Return a safe explanation
- Ask the user to remove sensitive content and retry
6) Add allowlists and context-based controls
Safety is better when it’s contextual:
- Internal staff may be allowed to access more data than external users
- Certain tenants may permit code generation, others may not
- Some endpoints may allow summarization of documents, but not extraction of personal data
7) Log carefully
Log enough for auditing, but avoid storing sensitive raw text:
- Store policy decision, reason code, model used, user/tenant ID
- Hash or redact prompt/output content
- Keep trace IDs for incident investigation
- Retain only as long as necessary
8) Example gateway flow
A practical pipeline:
- Receive prompt
- Authenticate user and identify tenant/role
- Run secret/PII scan
- Run prompt-injection and safety classifier
- Apply policy decision:
- allow
- redact
- block
- route
- Send sanitized prompt to model
- Receive output
- Run output moderation and leakage detection
- Redact or block as needed
- Log outcome and metrics
9) Suggested policy tiers
You can make this manageable by defining levels:
- Tier 0: allow
- Tier 1: redact sensitive data
- Tier 2: warn or require confirmation
- Tier 3: block and return safe refusal
- Tier 4: escalate to human review
10) Test with adversarial cases
Create a test set for:
- Prompt injection
- Obfuscated secrets
- Multilingual abuse
- Unicode trickery
- Nested instructions
- RAG document injections
- Model output echoing secrets
- Long-context leakage
11) If you’re implementing this technically
Common gateway components:
- Middleware for request/response interception
- Policy engine like OPA or custom rules
- Moderation API or internal classifier service
- DLP/PII detector
- Structured logging
- Feature flags for gradual rollout
12) Minimal pseudo-flow
request -> auth -> input scan -> policy check -> sanitize -> LLM call -> output scan -> sanitize/block -> response
13) Best practice summary
- Filter both input and output
- Combine regex, classifiers, and rules
- Assume all user and retrieved content is untrusted
- Redact when possible, block when necessary
- Log safely and monitor false positives/negatives
- Continuously red-team your gateway
If you want, I can also provide:
- a reference architecture for an AI gateway,
- sample policy rules, or
- code examples in Python/Node/Go.
Brands