Clinical Benchmarks

HealthBench: current results

OpenAI · 5,000 conversations · index updated August 16, 2026

Claude Opus 5 holds the top current result on HealthBench, 67.1 as of 2026-08, per OpenAI Deployment Safety Hub (GPT-5.6 system card + August 2026 updates); benchlm.ai and llm-stats.com mirror. 5,000 realistic multi-turn health conversations graded against physician-written rubrics (48,562 criteria) covering accuracy, completeness, context awareness, communication, and instruction following. OpenAI now also reports a length-adjusted variant that penalizes verbosity.

Current results

Result detail

#modelscoreas of
1Anthropic logoClaude Opus 5 Anthropic
raw/unadjusted, Anthropic system-card protocol (max effort, no tools, five trials); length-adjusted 57.8; mirrored on benchlm.ai
67.12026-08
2BBaichuan-M3 Baichuan
self-run in Baichuan-M3 paper (arXiv 2602.06570)
65.12026-02
3OpenAI logoGPT-5.2-High OpenAI
as run by Baichuan in the M3 paper, not OpenAI-reported
63.32026-02
4OpenAI logoGPT-5.6 Sol OpenAI
length-adjusted, max reasoning effort (55.6 unadjusted), GPT-5.6 system card 2026-07-09
57.02026-06
5OpenAI logoGPT-5.6 Terra OpenAI
length-adjusted (58.7 unadjusted), max reasoning effort
57.02026-06
6OpenAI logoGPT-5.5 OpenAI
length-adjusted (58.4 unadjusted), comparison row in GPT-5.6 system card
56.52026-06
7OpenAI logoGPT-5.6 Luna OpenAI
length-adjusted (55.4 unadjusted), max reasoning effort
55.82026-06
8OpenAI logoGPT-5.6 Sol (August) OpenAI
ChatGPT production/Instant deployment setting, length-adjusted (52.1 unadjusted), GPT-5.6 August Updates PDF 2026-08-06
55.02026-08

Scores appear exactly as OpenAI Deployment Safety Hub (GPT-5.6 system card + August 2026 updates); benchlm.ai and llm-stats.com mirror publishes them (mixed sources). Cross-source rows are not directly comparable: Anthropic reports a raw score under its own protocol, OpenAI leads with length-adjusted numbers at maximum reasoning effort and reports production Instant settings separately, and Baichuan ran its competitors itself. The benchmark's official grader is GPT-4.1. OpenAI has said the parent set is approaching a noise ceiling for frontier models and points to HealthBench Professional for continued measurement.

About the benchmark

publisherOpenAI
categoryrubric-graded benchmarks
released2025-05
size5,000 conversations
scale0-100 rubric-point percentage (some sites display 0-1), higher better; length-adjusted and unadjusted variants
result basismixed sources
sourceOpenAI Deployment Safety Hub (GPT-5.6 system card + August 2026 updates); benchlm.ai and llm-stats.com mirror
last frontier result2026-08

What is HealthBench?

HealthBench is a rubric-graded benchmark from OpenAI, released 2025-05: 5,000 conversations, scored on a 0-100 rubric-point percentage (some sites display 0-1) scale. 5,000 realistic multi-turn health conversations graded against physician-written rubrics (48,562 criteria) covering accuracy, completeness, context awareness, communication, and instruction following. OpenAI now also reports a length-adjusted variant that penalizes verbosity.

Which model leads HealthBench?

Claude Opus 5 (Anthropic) holds the top current result on HealthBench at 67.1, per OpenAI Deployment Safety Hub (GPT-5.6 system card + August 2026 updates); benchlm.ai and llm-stats.com mirror, as of 2026-08.

Where do the HealthBench numbers come from?

From OpenAI Deployment Safety Hub (GPT-5.6 system card + August 2026 updates); benchlm.ai and llm-stats.com mirror (mixed sources). Cross-source rows are not directly comparable: Anthropic reports a raw score under its own protocol, OpenAI leads with length-adjusted numbers at maximum reasoning effort and reports production Instant settings separately, and Baichuan ran its competitors itself. The benchmark's official grader is GPT-4.1. OpenAI has said the parent set is approaching a noise ceiling for frontier models and points to HealthBench Professional for continued measurement.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.