Prompt
I want a production monitoring setup that catches drift, hallucinations, and policy issues after deploy
Latest observation
Here’s a practical production monitoring setup for catching drift, hallucinations, and policy issues after deployment. I’ll frame it as a system you can actually run, not just a list of ideas.
1) What to monitor
A. Drift
Track whether production traffic is changing relative to your training / validation data.
Input drift
- Token/feature distribution changes
- Topic shifts
- User segment shifts
- Prompt length / structure changes
- Retrieval context changes (if using RAG)
Output drift
- Answer length
- Refusal rate
- Citation rate
- Confidence/uncertainty proxies
- Intent distribution
- Tool usage distribution
Behavioral drift
- Latency changes
- Fallback rate
- Escalation rate
- Human review rate
- Success / completion rate
B. Hallucinations
You usually can’t “detect hallucination” perfectly, so monitor for signals that strongly correlate with it.
Useful signals
- Groundedness: is the answer supported by retrieved sources?
- Citation validity: do cited sources actually support the claim?
- Retrieval coverage: did the model use the right docs?
- Contradictions: answer conflicts with source material
- Unsupported factual claims
- Overconfident language on uncertain queries
- Tool-result mismatch: model claims a tool returned something it didn’t
For RAG or tool-using systems, this is much more tractable.
C. Policy issues
Detect violations of your safety/product policy.
Examples
- Disallowed content
- Privacy leakage
- PII exposure
- Harassment / hate / sexual content
- Medical/legal/financial advice beyond allowed scope
- Prompt injection success
- Unsafe tool calls
- Unauthorized actions
- Refusal failures
2) Core architecture
Use a three-layer monitoring stack:
Layer 1: Real-time guardrails
Catch bad outputs before they reach users.
- Input filters
- Prompt injection detection
- Output moderation
- PII redaction
- Tool permission checks
- Schema validation for structured outputs
This is your “block or rewrite” layer.
Layer 2: Online observability
Capture every request/response pair with metadata.
Log:
- Prompt
- System prompt version
- Model version
- Temperature / decoding settings
- Retrieved documents and scores
- Tool calls and results
- User segment / locale / channel
- Latency
- Tokens in/out
- Safety classifier scores
- User feedback
- Escalation / correction signals
Make logs reproducible and versioned.
Layer 3: Offline evaluators
Run asynchronous checks on sampled traffic.
Use:
- Rule-based checks
- LLM-as-judge with strict rubrics
- Retrieval-grounded factuality scoring
- Policy classifiers
- Human review for edge cases
This layer catches issues that are too expensive or ambiguous for real-time blocking.
3) Drift detection design
Data drift
Compare live traffic against a baseline.
Recommended metrics:
- PSI (Population Stability Index) for numeric/categorical distributions
- KL divergence / Jensen-Shannon divergence
- Embedding centroid shift
- Topic distribution change
- Prompt template entropy
Output drift
Track:
- Refusal rate
- Unsafe completion rate
- Average confidence score
- Hallucination score
- Citation coverage
- Tool-call rate
- Completion length
- Retry rate
Thresholding
Use:
- Static thresholds for critical metrics
- Dynamic thresholds based on rolling windows
- Segment-specific baselines
- Alert severity levels
Example:
- Warning if JS divergence > 0.1 for 3 hours
- Critical if hallucination score rises 2x above baseline and user complaints also rise
4) Hallucination monitoring approach
If you use RAG
Best practice is to evaluate groundedness per answer.
Check:
- Did retrieval return relevant sources?
- Are claims in the answer supported by sources?
- Did the model cite the right passage?
- Does the answer include unsupported claims?
Metrics:
- Retrieval precision@k
- Answer support rate
- Citation precision
- Unsupported claim count
- Source contradiction rate
If you do not use RAG
Use proxy evaluators:
- Consistency checks across multiple generations
- Self-consistency
- Uncertainty heuristics
- User correction rate
- Expert review sampling
Practical guardrail
Have the model:
- cite evidence when possible
- say “I’m not sure” when evidence is missing
- abstain rather than invent
5) Policy monitoring
Build a policy taxonomy and map each rule to a detector.
Common detectors
- Classifiers for unsafe categories
- Regex/rule checks for PII and secrets
- Prompt injection signatures
- Output schema validators
- Tool authorization policy engine
- Escalation classifier
Policy event schema
For each violation log:
- policy_id
- severity
- detection method
- evidence
- model/version
- user/session
- action taken
- reviewer disposition
Action matrix
- Low severity: log only
- Medium severity: warn / redact / ask for clarification
- High severity: block response
- Critical: block + escalate + alert on-call
6) Human-in-the-loop review
You need human review for:
- low-confidence hallucination cases
- policy edge cases
- samples from new traffic segments
- newly released prompt/model versions
- high-risk domains
Sampling strategy:
- 100% review for critical alerts
- Stratified random sample of normal traffic
- Oversample:
- long prompts
- low retrieval confidence
- tool calls
- policy classifier borderline scores
- negative user feedback
7) Alerting and dashboards
Dashboards
Show by model version, prompt version, traffic segment:
- Drift metrics
- Hallucination score
- Safety violation rates
- Refusal rate
- Escalation rate
- Latency
- Cost
- User satisfaction
Alerts
Trigger on:
- sharp changes vs baseline
- combined signals, not single metrics
- sustained degradation
- critical policy events
Example alert: “RAG groundedness dropped 18% over baseline in finance queries after model v4 rollout; unsafe advice rate up 3x.”
8) Recommended implementation stack
A common production stack looks like:
- Telemetry: OpenTelemetry / custom event logging
- Storage: warehouse + object store
- Streaming: Kafka / PubSub / Kinesis
- Dashboards: Grafana / Superset / Looker
- Moderation: policy classifiers + vendor moderation APIs
- LLM evals: rubric-based evaluators, regression suites
- Human review: internal labeling tool
- Feature store / embedding store: for drift comparisons
- Alerting: PagerDuty / Slack / email
9) Rollout strategy
Before launch
- Build golden eval set
- Create policy test suite
- Record baseline metrics
- Test prompt injection and abuse cases
- Establish rollback criteria
After launch
- Shadow mode first, if possible
- Gradual traffic ramp
- Compare canary vs control
- Tight alerts for new versions
- Weekly review of samples and failures
10) Minimum viable setup
If you want the simplest useful version, do this:
- Log every request/response with model/prompt/version metadata
- Add moderation + PII filtering on inputs/outputs
- For RAG, log retrieved docs and run groundedness checks
- Run daily offline sampling with an LLM judge and human review
- Track drift on input embeddings, refusal rate, and output length
- Alert on policy violations, retrieval failure, and sharp metric shifts
- Keep a rollback path for prompt/model changes
11) Practical metric bundle to start with
If you need a starter dashboard, track:
Traffic
- request volume
- unique users
- segment mix
Quality
- user thumbs up/down
- task completion rate
- hallucination score
- groundedness score
Safety
- policy violation rate
- PII leak rate
- refusal rate
- escalation rate
Drift
- input embedding shift
- topic distribution change
- output length change
- retrieval relevance change
Ops
- latency p50/p95
- error rate
- cost per request
- tool failure rate
12) If you want, I can help you turn this into:
- a concrete architecture diagram,
- a metrics spec,
- a database/event schema,
- or an implementation plan in Python/SQL/Kafka.
If you tell me whether your system is RAG, tool-using agent, or plain chat, I can tailor the monitoring design precisely.