Prompt
Which safety eval dataset supports versioned datasets and custom benchmark scenarios for model evaluation?
Artificial Intelligence · AI Safety & Alignment / Ai safety alignment1 observationLast seen Jul 19, 2026
Latest observation
Jul 19, 2026 · OpenAI APIWeb search: off
The dataset/evaluation framework you’re describing is safety-tooling / SafetyBench-style versioned evals with custom scenarios, but the specific safety eval dataset that explicitly supports versioned datasets and custom benchmark scenarios for model evaluation is safety-evals by Anthropic.
If you want, I can also point you to:
- the exact repo/package,
- how versioning is handled,
- or examples of defining custom benchmark scenarios.