HealthBench Hard: current results
OpenAI · 1,000 conversations · index updated August 16, 2026
Muse Spark holds the top current result on HealthBench Hard, 0.428 as of 2026-08, per healthbenchhard.ai. The bottom fifth of HealthBench: 1,000 conversations where frontier models failed most at the May 2025 release, still graded on the original physician-written rubrics.
This benchmark has a dedicated full leaderboard, with methodology and per-model pages, at healthbenchhard.ai. The top of its table is mirrored below.
Current top results
- 1
Muse Spark0.428
- 2
GPT-5.6 Sol0.331
- 3
GPT-5.6 Terra0.327
Result detail
| # | model | score | as of | |
|---|---|---|---|---|
| 1 | Muse Spark Meta | 0.428 | 2026-08 | |
| 2 | GPT-5.6 Sol OpenAI | 0.331 | 2026-08 | |
| 3 | GPT-5.6 Terra OpenAI | 0.327 | 2026-08 | |
Scores appear exactly as healthbenchhard.ai publishes them (mixed sources). Default grader GPT-4.1. Best score at the May 2025 release was o3's 0.320.
About the benchmark
| publisher | OpenAI |
|---|---|
| category | rubric-graded benchmarks |
| released | 2025-05 |
| size | 1,000 conversations |
| scale | 0 to 1, higher is better |
| result basis | mixed sources |
| source | healthbenchhard.ai |
| last frontier result | 2026-08 |
What is HealthBench Hard?
HealthBench Hard is a rubric-graded benchmark from OpenAI, released 2025-05: 1,000 conversations, scored on a 0 to 1 scale. The bottom fifth of HealthBench: 1,000 conversations where frontier models failed most at the May 2025 release, still graded on the original physician-written rubrics.
Which model leads HealthBench Hard?
Muse Spark (Meta) holds the top current result on HealthBench Hard at 0.428, per healthbenchhard.ai, as of 2026-08.
Where do the HealthBench Hard numbers come from?
From healthbenchhard.ai (mixed sources). Default grader GPT-4.1. Best score at the May 2025 release was o3's 0.320.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.