Clinical Benchmarks

WHBench: current results

Independent researchers (Maurya, Govindgari, Kumar) · 47 scenarios / 3,100 scored responses across 22 models · index updated August 16, 2026

Claude Opus 4.6 holds the top current result on WHBench, 72.1% as of 2026-07, per arXiv paper (v2 revised 2026-07-23). Women's health: 47 expert-crafted scenarios across 10 topics graded on a 23-criterion rubric for clinical accuracy, safety, equity, and guideline adherence; targets failure modes like outdated guidelines, unsafe omissions, dosing errors, equity blind spots.

Current results

Result detail

#modelscoreas of
1Anthropic logoClaude Opus 4.6 Anthropic
95% CI 69.6-74.4; evaluations run March 2026
72.1%2026-07
2Anthropic logoClaude Sonnet 4.6 Anthropic
95% CI 64.5-69.6
67.1%2026-07
3OpenAI logoGPT-5.4 OpenAI
95% CI 64.5-69.2
66.8%2026-07
4Google logoGemini 3 Flash Preview Google64.7%2026-07
5OpenAI logoGPT-4.1 OpenAI51.8%2026-07
6OpenAI logoGPT-4o OpenAI44.6%2026-07

Scores appear exactly as arXiv paper (v2 revised 2026-07-23) publishes them (independently run). An academic study with expert validation rather than a live leaderboard; the model set was frozen in March 2026, before GPT-5.6 and the Claude 5 family shipped.

About the benchmark

publisherIndependent researchers (Maurya, Govindgari, Kumar)
categoryrubric-graded benchmarks
released2026-04
size47 scenarios / 3,100 scored responses across 22 models
scalemean normalized percentage 0-100, higher better
result basisindependently run
sourcearXiv paper (v2 revised 2026-07-23)
last frontier result2026-07

What is WHBench?

WHBench is a rubric-graded benchmark from academic team, released 2026-04: 47 scenarios / 3,100 scored responses across 22 models, scored on a mean normalized percentage 0-100 scale. Women's health: 47 expert-crafted scenarios across 10 topics graded on a 23-criterion rubric for clinical accuracy, safety, equity, and guideline adherence; targets failure modes like outdated guidelines, unsafe omissions, dosing errors, equity blind spots.

Which model leads WHBench?

Claude Opus 4.6 (Anthropic) holds the top current result on WHBench at 72.1%, per arXiv paper (v2 revised 2026-07-23), as of 2026-07.

Where do the WHBench numbers come from?

From arXiv paper (v2 revised 2026-07-23) (independently run). An academic study with expert validation rather than a live leaderboard; the model set was frozen in March 2026, before GPT-5.6 and the Claude 5 family shipped.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.