Prompt
I need a way to compare model outputs on my own labeled dataset
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
A good way to compare model outputs on your own labeled dataset is to set up a small evaluation harness that runs each model on the same inputs, then scores their outputs against your labels with metrics that fit your task.
Basic approach
-
Prepare your dataset
- Each example should include:
input/ promptgold_label/ reference answer- optional metadata like category, difficulty, or source
- Each example should include:
-
Run each model on the same inputs
- Save each model’s prediction for every example
- Keep outputs in a consistent format
-
Define evaluation metrics
- For classification: accuracy, precision, recall, F1
- For extraction: exact match, token F1
- For ranking or scoring: Spearman/Pearson correlation
- For generated text: exact match when possible, otherwise ROUGE/BLEU or human review
- For LLM-style judging: pairwise preference or rubric-based scoring
-
Compare results
- Overall score per model
- Scores by slice/category
- Error analysis on disagreements
-
Use statistical testing if needed
- Bootstrap confidence intervals
- McNemar’s test for paired classification
- Paired t-test or bootstrap for continuous metrics
Simple evaluation schema
You can store your data like this:
[
{
"id": 1,
"input": "What is 2+2?",
"label": "4"
},
{
"id": 2,
"input": "Name a mammal that can fly.",
"label": "bat"
}
]
And predictions like this:
[
{
"id": 1,
"model_a": "4",
"model_b": "four"
},
{
"id": 2,
"model_a": "bat",
"model_b": "bird"
}
]
Example comparison logic
For a simple exact-match task:
from sklearn.metrics import accuracy_score
y_true = ["4", "bat"]
y_pred_a = ["4", "bat"]
y_pred_b = ["four", "bird"]
acc_a = accuracy_score(y_true, y_pred_a)
acc_b = accuracy_score(y_true, y_pred_b)
print("Model A:", acc_a)
print("Model B:", acc_b)
If your outputs are free-form text
Exact match is often too strict. You can:
- normalize text: lowercase, strip punctuation, remove extra whitespace
- use token-based F1
- use semantic similarity embeddings
- use an LLM as a judge with a fixed rubric
Example normalization:
import re
def normalize(text):
text = text.lower().strip()
text = re.sub(r"\W+", " ", text)
return " ".join(text.split())
Recommended tooling
Depending on your use case:
- Python + pandas for a lightweight custom setup
- scikit-learn for standard metrics
- evaluate by Hugging Face for NLP metrics
- OpenAI Evals or lm-eval-harness if you want a more structured benchmark setup
- Weights & Biases or MLflow for tracking runs
A practical workflow
- Export your labeled set to CSV/JSON
- Generate outputs from each model
- Compute metrics in Python
- Slice performance by category
- Manually inspect disagreements
- Repeat after prompt/model changes
If you want, I can help you with one of these next:
- a Python script to compare two models on a labeled CSV
- a template dataset format
- a pairwise evaluation setup
- a metric recommendation based on your task type