Clinical Benchmarks

Rubric-graded benchmarks

5 tracked · updated August 16, 2026

These benchmarks grade free-text answers against rubrics that physicians or domain experts wrote, criterion by criterion, instead of scoring a picked letter. They reward completeness, accuracy, and safe framing at once, which is why frontier scores on them run low and move slowly. All of them descend from the idea that a health answer is a judgment to be audited, not a fact to be matched.

HealthBench Professional

OpenAI · 525 tasks
  1. 1Anthropic logoClaude Fable 50.660
  2. 2OpenAI logoGPT-5.6 Sol0.605
  3. 3Anthropic logoClaude Opus 50.598
via healthbenchprofessional.com · updated 2026-08full detail

HealthBench Hard

OpenAI · 1,000 conversations
  1. 1Meta logoMuse Spark0.428
  2. 2OpenAI logoGPT-5.6 Sol0.331
  3. 3OpenAI logoGPT-5.6 Terra0.327

HealthBench

OpenAI · 5,000 conversations
  1. 1Anthropic logoClaude Opus 567.1
  2. 2BBaichuan-M365.1
  3. 3OpenAI logoGPT-5.2-High63.3
  4. 4OpenAI logoGPT-5.6 Sol57.0
  5. 5OpenAI logoGPT-5.6 Terra57.0
via OpenAI Deployment Safety Hub · updated 2026-08full detail

Health Optimization Bench

healthoptimizationbench.com · 89 tasks
  1. 1Anthropic logoClaude Fable 583.8
  2. 2xAI logoGrok 4.681.3
  3. 3Anthropic logoClaude Opus 578.3
via healthoptimizationbench.com · updated 2026-08full detail

WHBench

academic team · 47 scenarios
  1. 1Anthropic logoClaude Opus 4.672.1%
  2. 2Anthropic logoClaude Sonnet 4.667.1%
  3. 3OpenAI logoGPT-5.466.8%
  4. 4Google logoGemini 3 Flash Preview64.7%
  5. 5OpenAI logoGPT-4.151.8%

Which rubric-graded benchmarks have current frontier-model results?

5 as of August 16, 2026: HealthBench Professional (Claude Fable 5 leads at 0.660); HealthBench Hard (Muse Spark leads at 0.428); HealthBench (Claude Opus 5 leads at 67.1); Health Optimization Bench (Claude Fable 5 leads at 83.8); WHBench (Claude Opus 4.6 leads at 72.1%).

The other categories sit on the index: agentic and workflow benchmarks, documentation and coding benchmarks, safety benchmarks, knowledge and exam benchmarks, composite indices.