Prompt
How do I log every AI request and response for production auditing?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
To log every AI request/response for production auditing, the safest pattern is:
- Put logging at the API boundary
- Capture both the prompt and the model output
- Store metadata needed for auditability
- Redact/secure sensitive data
- Make logs immutable and searchable
What to log
For each AI call, record:
- Request ID / correlation ID
- Timestamp
- User ID / tenant ID
- Model name / version
- Full request payload
- system prompt
- developer prompt
- user prompt
- tool/function definitions
- parameters like temperature, max tokens
- Full response payload
- assistant message
- tool calls
- finish reason
- Usage metadata
- input/output tokens
- latency
- cost if available
- Error details if the call fails
- Safety/moderation flags if you use them
Important cautions
1) Don’t log secrets or sensitive PII in plaintext
If prompts may contain:
- passwords
- API keys
- SSNs
- health data
- payment data
then redact before storing.
2) Separate audit logs from application logs
Audit logs should be:
- append-only
- access-controlled
- retained according to policy
- tamper-evident if possible
3) Consider storing hashes for integrity
A hash of the request/response can prove the record wasn’t altered later.
Recommended architecture
App → AI wrapper/service → logger → AI provider
Instead of calling the provider directly throughout your code, route all calls through a single wrapper that:
- assigns request ID
- logs request before sending
- logs response after receiving
- handles redaction
- writes to your audit store
Example: Python wrapper
import time
import json
import uuid
import logging
from datetime import datetime, timezone
logger = logging.getLogger("ai_audit")
logger.setLevel(logging.INFO)
def redact(text: str) -> str:
# Replace with real redaction logic
return text.replace("sk-", "[REDACTED_KEY]")
def log_audit(event_type, data):
logger.info(json.dumps({
"event_type": event_type,
"timestamp": datetime.now(timezone.utc).isoformat(),
**data
}))
def call_ai(client, messages, model="gpt-4.1", **kwargs):
request_id = str(uuid.uuid4())
start = time.time()
safe_messages = [
{**m, "content": redact(m.get("content", ""))}
for m in messages
]
log_audit("ai_request", {
"request_id": request_id,
"model": model,
"messages": safe_messages,
"params": kwargs
})
try:
response = client.chat.completions.create(
model=model,
messages=messages,
**kwargs
)
duration_ms = int((time.time() - start) * 1000)
log_audit("ai_response", {
"request_id": request_id,
"model": model,
"duration_ms": duration_ms,
"response": response.model_dump() if hasattr(response, "model_dump") else str(response),
})
return response
except Exception as e:
duration_ms = int((time.time() - start) * 1000)
log_audit("ai_error", {
"request_id": request_id,
"model": model,
"duration_ms": duration_ms,
"error": str(e)
})
raise
Example: what a log record might look like
{
"event_type": "ai_request",
"timestamp": "2026-10-10T12:00:00Z",
"request_id": "c1d2e3f4",
"user_id": "12345",
"tenant_id": "acme",
"model": "gpt-4.1",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Summarize the attached contract."}
],
"params": {
"temperature": 0.2,
"max_tokens": 500
}
}
Production best practices
- Use structured logging (JSON)
- Send logs to a centralized system like:
- Elasticsearch / OpenSearch
- Splunk
- Datadog
- CloudWatch / Stackdriver / Azure Monitor
- Encrypt logs at rest and in transit
- Set role-based access
- Define retention policies
- Mask/redact sensitive fields
- Log tool calls separately if the model can invoke tools
- Capture prompt versioning so you know exactly what instruction set was used
If you need full audit traceability
Store:
- raw request
- raw response
- normalized view
- hashes
- user/context metadata
- model/version metadata
- tool execution details
And optionally write to:
- database for queryability
- object storage for immutable archive
- SIEM for monitoring
If you want, I can provide a ready-to-drop-in example for:
- Python / FastAPI
- Node.js / Express
- OpenAI SDK wrapper
- LangChain/LlamaIndex logging
- AWS/GCP/Azure production setup