Prompt
How do I enforce prompt and output policies before requests hit model providers?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
To enforce prompt and output policies before requests reach model providers, put a policy enforcement layer in front of your LLM calls. Think of it as a gateway that inspects, rewrites, blocks, or routes requests/responses before anything is sent onward.
Recommended architecture
Client → Policy Gateway → Model Provider
- Input policy checks on the prompt/messages
- Prompt transformation if needed
- Model routing only after passing policy
- Output policy checks on the provider response
- Block/redact/repair before returning to the user
What to enforce
1) Prompt-side policies
Inspect the request before forwarding it:
- PII/PHI/secrets detection
- Prompt injection / jailbreak heuristics
- Disallowed content categories
- Tenant/data-boundary validation
- Rate limits and quota checks
- Allowed model / region / tool permissions
Actions:
- Allow
- Deny
- Redact
- Transform (e.g., strip secrets, normalize, add safety system prompt)
- Route to a safer model or sandbox
2) Output-side policies
Inspect model responses before returning them:
- Toxicity / harassment
- Disallowed instructions
- Sensitive data leakage
- Policy-specific formatting constraints
- Citation requirements / hallucination checks (if applicable)
Actions:
- Block
- Redact
- Regenerate
- Fallback to a safe response template
- Escalate to human review
Implementation pattern
A. Use a policy engine
Common choices:
- OPA (Open Policy Agent)
- Cedar
- Custom rules engine
- Regex + detectors for simpler setups
Policies should be versioned and auditable.
B. Add pre-request middleware
In your API service:
- Receive request
- Normalize input
- Run policy checks
- Decide action
- Only then call provider
Pseudo-flow:
if violates_input_policy(request):
return blocked_response
request = sanitize(request)
response = call_model_provider(request)
if violates_output_policy(response):
return blocked_or_redacted_response
return response
C. Keep policy decisions separate from model logic
Do not rely on the model to police itself for critical controls. Enforce in your application layer.
Example: simple middleware
def handle_llm_request(req):
decision = policy_engine.check_input(req)
if decision.action == "block":
return {"error": "Request blocked by policy"}
if decision.action == "redact":
req = redact_sensitive_fields(req)
provider_resp = call_provider(req)
out_decision = policy_engine.check_output(provider_resp)
if out_decision.action == "block":
return {"error": "Response blocked by policy"}
if out_decision.action == "redact":
provider_resp = redact(provider_resp)
return provider_resp
Best practices
- Fail closed for sensitive environments
- Log decisions with reason codes
- Version policies and test them
- Use allowlists for models, tools, and destinations
- Separate tenants and contexts
- Redact before logging
- Add a human review path for borderline cases
If you use a gateway/proxy
You can enforce policies via:
- API gateway plugins
- LLM proxy services
- Service mesh filters
- Reverse proxy middleware
This is ideal if multiple apps/services call models.
Common pitfalls
- Checking only output and ignoring prompt leaks
- Letting prompts reach the provider before scanning
- Overblocking with fragile regexes
- Not handling streaming responses incrementally
- Not validating tool/function calls separately
Practical recommendation
If you’re starting today:
- Put a middleware/gateway in front of model APIs
- Use a policy engine for deterministic rules
- Add PII and secret detection
- Add prompt injection and content moderation checks
- Validate output before returning it
- Log every policy decision for audit
If you want, I can sketch a reference design for:
- Python/FastAPI
- Node/Express
- OpenAI-compatible proxy
- OPA/Cedar-based policy enforcement