Prompt

I need a way to compare model outputs on my own labeled dataset

Artificial Intelligence / AI Developer Tools1 observationLast seen Oct 9, 2026

Latest observation

Oct 9, 2026OpenAI APIWeb search: off

A good way to compare model outputs on your own labeled dataset is to set up a small evaluation harness that runs each model on the same inputs, then scores their outputs against your labels with metrics that fit your task.

Basic approach

  1. Prepare your dataset

    • Each example should include:
      • input / prompt
      • gold_label / reference answer
      • optional metadata like category, difficulty, or source
  2. Run each model on the same inputs

    • Save each model’s prediction for every example
    • Keep outputs in a consistent format
  3. Define evaluation metrics

    • For classification: accuracy, precision, recall, F1
    • For extraction: exact match, token F1
    • For ranking or scoring: Spearman/Pearson correlation
    • For generated text: exact match when possible, otherwise ROUGE/BLEU or human review
    • For LLM-style judging: pairwise preference or rubric-based scoring
  4. Compare results

    • Overall score per model
    • Scores by slice/category
    • Error analysis on disagreements
  5. Use statistical testing if needed

    • Bootstrap confidence intervals
    • McNemar’s test for paired classification
    • Paired t-test or bootstrap for continuous metrics

Simple evaluation schema

You can store your data like this:

[
  {
    "id": 1,
    "input": "What is 2+2?",
    "label": "4"
  },
  {
    "id": 2,
    "input": "Name a mammal that can fly.",
    "label": "bat"
  }
]

And predictions like this:

[
  {
    "id": 1,
    "model_a": "4",
    "model_b": "four"
  },
  {
    "id": 2,
    "model_a": "bat",
    "model_b": "bird"
  }
]

Example comparison logic

For a simple exact-match task:

from sklearn.metrics import accuracy_score

y_true = ["4", "bat"]
y_pred_a = ["4", "bat"]
y_pred_b = ["four", "bird"]

acc_a = accuracy_score(y_true, y_pred_a)
acc_b = accuracy_score(y_true, y_pred_b)

print("Model A:", acc_a)
print("Model B:", acc_b)

If your outputs are free-form text

Exact match is often too strict. You can:

  • normalize text: lowercase, strip punctuation, remove extra whitespace
  • use token-based F1
  • use semantic similarity embeddings
  • use an LLM as a judge with a fixed rubric

Example normalization:

import re

def normalize(text):
    text = text.lower().strip()
    text = re.sub(r"\W+", " ", text)
    return " ".join(text.split())

Recommended tooling

Depending on your use case:

  • Python + pandas for a lightweight custom setup
  • scikit-learn for standard metrics
  • evaluate by Hugging Face for NLP metrics
  • OpenAI Evals or lm-eval-harness if you want a more structured benchmark setup
  • Weights & Biases or MLflow for tracking runs

A practical workflow

  1. Export your labeled set to CSV/JSON
  2. Generate outputs from each model
  3. Compute metrics in Python
  4. Slice performance by category
  5. Manually inspect disagreements
  6. Repeat after prompt/model changes

If you want, I can help you with one of these next:

  • a Python script to compare two models on a labeled CSV
  • a template dataset format
  • a pairwise evaluation setup
  • a metric recommendation based on your task type

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.