Company
Gpqa
Sample prompts where it appears
What's the most effective safety eval suite for scoring model helpfulness versus harmlessness during release reviews?
Brands:Harmbench,Strongreject,Realtoxicityprompts,Advbench,Jailbreakbench
Can you recommend a benchmark suite for comparative model testing and safety regression coverage?
Brands:Helm,Lmsys,Chatbot Arena,Mmlu,Gsm8k
Unlock deeper brand history
Sign up free to see richer competitive landscapes, historical mention trends, and related model comparisons.
How did Obsurfable measure this?
Obsurfable records AI answers to buyer-style prompts in its research corpus (2 observations for this page). Metrics are distributions over observations, not a single static ranking.
Which AI systems does Obsurfable collect answers from?
OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, and Meta AI.