Prompt
Can you suggest a workflow for calibrating LLM-as-judge scores against human ratings over time?
Latest observation
Yes — a good workflow is to treat the LLM judge as a measurable scoring system that needs periodic alignment with human labels, drift monitoring, and recalibration. Here’s a practical workflow you can use.
1) Define the scoring target clearly
Before calibration, make sure both humans and the LLM judge are scoring the same thing.
- Task definition: e.g. helpfulness, factuality, tone, safety, completeness
- Rating scale: binary, 1–5 Likert, pairwise preference, etc.
- Rubric: explicit criteria and anchor examples for each score
- Aggregation rule: how multiple human ratings combine into the “gold” label
If the rubric is ambiguous, calibration will be unstable over time.
2) Build a fixed calibration set
Create a “gold” evaluation set that is rated by humans.
Recommended:
- Sample examples representative of real traffic
- Include edge cases and hard negatives
- Use multiple human raters per example
- Measure inter-rater agreement
- Resolve disagreements with adjudication or majority vote
Keep one version of this set stable so you can compare model behavior over time.
You may also want:
- Static calibration set for longitudinal tracking
- Rolling freshness set for new product/content distributions
3) Collect paired human and LLM scores
For each item in the calibration set:
- Human score(s)
- LLM judge score
- Optional metadata:
- task type
- prompt version
- model version
- language
- domain
- source channel
- time period
This lets you calibrate not just globally, but by segment.
4) Measure raw agreement and bias
Before adjusting anything, compute:
- Correlation
- Pearson for linear relation
- Spearman/Kendall for rank agreement
- Classification metrics if using thresholds
- Accuracy, F1, precision/recall
- Calibration metrics
- Expected calibration error (ECE) if scores map to probabilities
- Error analysis
- Mean signed error
- Over/under-scoring by category
- Confusion matrix for ordinal/binary labels
This tells you whether the LLM judge is systematically lenient, harsh, or inconsistent.
5) Fit a calibration transform
Map LLM scores to human-aligned scores using a simple transform first.
Common approaches:
If scores are numeric/ordinal
- Linear regression
- Isotonic regression for monotonic but nonlinear mapping
- Ordinal logistic regression
If scores are probabilities
- Platt scaling
- Temperature scaling
- Beta calibration
If outputs are pairwise preferences
- Convert pairwise win rates to a latent score scale, then calibrate to human ratings.
A simple, monotonic transform is often best because it preserves ranking while correcting scale bias.
6) Validate on a held-out set
Don’t calibrate and evaluate on the same examples only.
Use:
- Train calibration mapping on one subset
- Validate on another subset
- Track:
- correlation improvement
- reduced bias
- better threshold performance
- stability across segments
If a calibration mapping helps only on the train subset, it’s overfitting.
7) Deploy with versioning
Treat the calibration layer as a versioned artifact:
- LLM judge model version
- Prompt version
- Calibration mapping version
- Human rubric version
- Data window used to fit it
This is important because any change in prompt or judge model can invalidate old calibration.
8) Monitor drift over time
Human preferences and content distributions change, so calibration is not one-and-done.
Monitor regularly:
- Score distribution shifts
- Agreement metrics over time
- Residuals by segment
- Drift in domain mix
- Changes in human rater behavior
Set triggers for recalibration when:
- correlation drops
- bias exceeds a threshold
- content mix changes significantly
- new policy/rubric is introduced
- judge model or prompt changes
9) Use a rolling recalibration schedule
A good operational pattern:
- Daily/weekly: monitor judge-human gap on recent samples
- Monthly/quarterly: refit calibration mapping if drift is detected
- Immediately: recalibrate after judge model updates or rubric changes
A rolling window is useful:
- Fit on last N weeks/months of human-labeled data
- Keep a fixed “historical benchmark” set for comparison
10) Add segment-level calibration
Global calibration often hides systematic errors.
Calibrate separately or at least report separately for:
- language
- topic/domain
- conversation length
- safety-sensitive vs general content
- user intent type
- prompt class
- output length
Sometimes a single global mapping is fine; sometimes you need a hierarchical model or per-segment offsets.
11) Use a human audit loop
Each cycle, sample:
- high-confidence judge positives
- high-confidence judge negatives
- borderline cases
- examples with largest judge-human disagreement
Use this for:
- identifying rubric ambiguities
- improving prompts
- detecting model failure modes
- updating human guidelines
This makes calibration an active quality-control loop, not just a statistical fit.
12) Maintain a decision policy separate from raw scores
If the LLM judge is used for decisions, define explicit thresholds after calibration.
Example:
- raw judge score → calibrated human-equivalent score
- then threshold:
- accept if ≥ 4.0
- review if 3.0–3.9
- reject if < 3.0
This is better than using raw LLM scores directly, because calibration makes thresholds more interpretable.
A simple end-to-end workflow
Here’s a compact version you can operationalize:
- Sample a representative set of items.
- Get human ratings using a stable rubric.
- Score with LLM judge using a fixed prompt/model version.
- Fit calibration mapping from LLM scores to human ratings.
- Evaluate on held-out data and by segment.
- Deploy mapping with versioning.
- Monitor drift weekly/monthly.
- Recalibrate when drift, rubric, or model changes.
- Audit disagreements to improve rubric and judge prompt.
Practical tips
- Prefer monotonic calibration if you want to preserve ranking.
- Keep a stable benchmark set that never changes.
- Use multiple human raters for noisy tasks.
- Track confidence intervals, not just point estimates.
- Don’t calibrate across fundamentally different tasks with one global mapping.
- Version everything: rubric, prompt, judge model, calibration set, transform.
If you want, I can also provide:
- a Python implementation sketch for calibration and monitoring, or
- a MLOps-style architecture for running this in production.