Prompt
I'm building a customer-facing classifier with an LLM and need repeatable evals
Latest observation
To get repeatable evals for a customer-facing LLM classifier, you want to treat evaluation like a small production system: fixed inputs, fixed prompts, fixed decoding, fixed metrics, and versioned artifacts.
1) Make the model behavior deterministic as much as possible
Use the most stable settings you can:
- temperature = 0
- top_p = 1
- fixed prompt
- fixed model version
- fixed system message
- If supported, seed the run
- Avoid tools / retrieval / web access during evals unless they are part of the production path
If the model still shows some nondeterminism, run each case multiple times and measure variance.
2) Use a locked eval set
Create a dataset that never changes unless explicitly versioned.
Include:
- Representative real customer examples
- Edge cases
- Ambiguous cases
- Adversarial examples
- Negative examples
- Class-balanced samples if possible
Best practice:
- Split into train/dev/test or at least dev vs. locked test
- Never tune on the locked test set
- Version the dataset:
eval_set_v1,eval_set_v2, etc.
3) Define the task output strictly
For classification, force the model to output only a schema you can parse.
Example:
{
"label": "billing_issue",
"confidence": 0.93
}
Or even just:
billing_issue
The tighter the output format, the more repeatable the eval.
4) Build a rubric with exact metrics
For a classifier, use standard metrics:
- Accuracy
- Precision / Recall / F1
- Confusion matrix
- Per-class metrics
- Coverage / abstention rate if the model can refuse
- Calibration if confidence matters
If class imbalance exists, don’t rely on accuracy alone. Use macro F1 and per-class recall.
5) Measure consistency, not just correctness
Repeatability matters in customer-facing systems.
Track:
- Same prompt, same input, same output rate
- Output variance across repeated runs
- Drift across model versions
- Stability under paraphrases
- Stability under minor formatting changes
A useful metric:
- Consistency score = percent of repeated runs that produce the same label
6) Add adversarial and paraphrase tests
For robustness, create test pairs like:
- Original text
- Paraphrased text
- Same meaning, different punctuation/casing
- Added noise or filler text
If a classifier changes label on small surface changes, that’s a red flag.
7) Separate “prompt evals” from “model evals”
When you change the prompt:
- Evaluate prompt version separately
- Keep model version constant
When you change the model:
- Re-run the same locked dataset
- Compare deltas by class and overall
This helps you know whether a regression came from the prompt or the model.
8) Log everything needed to reproduce
For each run, store:
- Model name/version
- Prompt version
- Decoding parameters
- Input dataset version
- Timestamp
- Output
- Raw response
- Parsed label
- Metrics script version
Use experiment tracking tools if helpful:
- MLflow
- Weights & Biases
- LangSmith
- OpenAI Evals-style harness
- A simple versioned CSV + Git can also work
9) Use a regression gate
Before deployment, compare against a baseline.
Example rules:
- Overall F1 must not drop by more than 1%
- No critical class recall may drop by more than 2%
- No increase in invalid output rate above 0.5%
This gives you a pass/fail decision instead of subjective review.
10) Human review for borderline cases
For customer-facing classifiers, have humans adjudicate:
- Ambiguous samples
- Disagreements between model and label
- High-confidence wrong predictions
This improves both eval quality and your label schema.
A practical eval loop
- Freeze dataset version
- Freeze prompt version
- Run model with deterministic settings
- Parse outputs
- Compute metrics
- Compare to baseline
- Review errors
- Update prompt/model
- Re-run the same locked test set
Minimal setup I’d recommend
If you want something simple and repeatable:
- JSONL eval set with gold labels
- Prompt version in Git
- temperature 0
- model version pinned
- script that outputs:
- accuracy
- macro F1
- per-class recall
- invalid JSON rate
- confusion matrix
- baseline comparison in CI
If you want, I can also give you:
- a ready-to-use eval harness template in Python,
- a dataset schema for classification evals, or
- a production checklist for customer-facing LLM classifiers.