Prompt

How do I add safety filters to prompts and outputs in an AI gateway?

Artificial Intelligence / AI Infrastructure1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

To add safety filters to both prompts and outputs in an AI gateway, use a layered approach: classify → enforce → log → review. A gateway is a good place to centralize this because it sits between your apps and model providers.

1) Define what you want to block or transform

Start with a clear policy for both inbound and outbound traffic:

  • Prompts (input)

    • PII: SSNs, credit cards, API keys, passwords
    • Prompt injection: “ignore previous instructions,” “reveal system prompt”
    • Harmful requests: malware, fraud, self-harm instructions, hate/harassment
    • Data exfiltration: requests to expose secrets, internal docs
  • Outputs

    • PII leakage
    • Unsafe advice or disallowed content
    • Policy violations: copyrighted text, confidential data, unsafe code
    • Hallucinated sensitive claims if your domain requires accuracy controls

2) Put filters in two places

Use both pre-processing and post-processing:

A. Input filtering before the model

Do this as close to the gateway edge as possible:

  • Pattern matching / regex for obvious secrets and PII
  • DLP/PII detection service
  • Prompt injection classifier
  • Topic moderation classifier
  • Context-aware allow/deny rules by user role, tenant, or endpoint

Actions:

  • Block high-risk requests
  • Redact sensitive parts
  • Warn / require confirmation
  • Route to a safer model or stricter policy tier

B. Output filtering after the model

Before returning the response:

  • Scan for PII/secrets
  • Run moderation/safety classification
  • Check for policy-specific forbidden content
  • Detect if the model echoed hidden instructions or internal data

Actions:

  • Replace risky spans with redaction
  • Refuse and return a safe completion
  • Truncate if necessary
  • Escalate to human review for borderline cases

3) Use multiple detection methods

No single filter is enough. Combine:

Deterministic checks

Good for:

  • API keys
  • JWTs
  • SSNs
  • Credit cards
  • Known secret patterns

ML-based classifiers

Good for:

  • Toxicity
  • Hate/harassment
  • Sexual content
  • Self-harm
  • Prompt injection
  • Sensitive intent

Policy/rule engine

Good for:

  • Tenant-specific rules
  • Role-based access control
  • Endpoint-specific restrictions
  • Region/compliance constraints

4) Protect against prompt injection

Treat user content as untrusted. For gateway design:

  • Separate system instructions from user input
  • Never let user content overwrite policy
  • Strip or flag meta-instructions like:
    • “ignore previous instructions”
    • “reveal your prompt”
    • “act as system”
  • If your app uses RAG, filter retrieved documents too, because they can carry injection text

Helpful practices:

  • Mark retrieved content as untrusted
  • Summarize or extract facts before sending to the model
  • Use tool/function calling with strict schemas
  • Limit what tools the model can access

5) Redact instead of raw block when possible

For usability, don’t always deny the request. Options:

  • Replace detected secrets with [REDACTED]
  • Mask PII partially
  • Return a safe explanation
  • Ask the user to remove sensitive content and retry

6) Add allowlists and context-based controls

Safety is better when it’s contextual:

  • Internal staff may be allowed to access more data than external users
  • Certain tenants may permit code generation, others may not
  • Some endpoints may allow summarization of documents, but not extraction of personal data

7) Log carefully

Log enough for auditing, but avoid storing sensitive raw text:

  • Store policy decision, reason code, model used, user/tenant ID
  • Hash or redact prompt/output content
  • Keep trace IDs for incident investigation
  • Retain only as long as necessary

8) Example gateway flow

A practical pipeline:

  1. Receive prompt
  2. Authenticate user and identify tenant/role
  3. Run secret/PII scan
  4. Run prompt-injection and safety classifier
  5. Apply policy decision:
    • allow
    • redact
    • block
    • route
  6. Send sanitized prompt to model
  7. Receive output
  8. Run output moderation and leakage detection
  9. Redact or block as needed
  10. Log outcome and metrics

9) Suggested policy tiers

You can make this manageable by defining levels:

  • Tier 0: allow
  • Tier 1: redact sensitive data
  • Tier 2: warn or require confirmation
  • Tier 3: block and return safe refusal
  • Tier 4: escalate to human review

10) Test with adversarial cases

Create a test set for:

  • Prompt injection
  • Obfuscated secrets
  • Multilingual abuse
  • Unicode trickery
  • Nested instructions
  • RAG document injections
  • Model output echoing secrets
  • Long-context leakage

11) If you’re implementing this technically

Common gateway components:

  • Middleware for request/response interception
  • Policy engine like OPA or custom rules
  • Moderation API or internal classifier service
  • DLP/PII detector
  • Structured logging
  • Feature flags for gradual rollout

12) Minimal pseudo-flow

request -> auth -> input scan -> policy check -> sanitize -> LLM call -> output scan -> sanitize/block -> response

13) Best practice summary

  • Filter both input and output
  • Combine regex, classifiers, and rules
  • Assume all user and retrieved content is untrusted
  • Redact when possible, block when necessary
  • Log safely and monitor false positives/negatives
  • Continuously red-team your gateway

If you want, I can also provide:

  1. a reference architecture for an AI gateway,
  2. sample policy rules, or
  3. code examples in Python/Node/Go.

Brands

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.