Healthcare AI Benchmarks: Current Frontier Results
Index updated August 16, 2026 · 17 benchmarks · 70 models
This index tracks every healthcare benchmark on which current frontier models hold a public result: 17 boards across 17 sources as of August 16, 2026, each score carrying the source it came from. No single model owns the field, though Claude Opus 5 leads more than one board. The table links to a page per benchmark and a page per model.
The benchmarks
One board per benchmark, bars scaled to that board's own leader. Scores sit on different scales and never compare across boards.
HealthBench Professional
OpenAI · 525 tasks- 1
Claude Fable 50.660
- 2
GPT-5.6 Sol0.605
- 3
Claude Opus 50.598
HealthBench Hard
OpenAI · 1,000 conversations- 1
Muse Spark0.428
- 2
GPT-5.6 Sol0.331
- 3
GPT-5.6 Terra0.327
HealthBench
OpenAI · 5,000 conversations- 1
Claude Opus 567.1
- 2BBaichuan-M365.1
- 3
GPT-5.2-High63.3
- 4
GPT-5.6 Sol57.0
- 5
GPT-5.6 Terra57.0
MAST (Medical AI Superintelligence Test)
ARISE AI Research Network · 6-benchmark composite- 1
GPT-5.6 Sol60.2%
- 2
Kimi K360.1%
- 3
Gemini 3.6 Flash59.3%
- 4
Gemini 3.1 Pro58.9%
- 5AQwen3.5 397B A17B57.9%
MedHELM
Stanford CRFM · 121 tasks- 1
Gemini 3.1 Pro (Preview)0.652
- 2
Gemini 3.5 Flash0.642
- 3
Muse Spark (2026-04-08)0.621
- 4
GPT-5.4 mini0.552
- 5
GPT-5.4 (2026-03-05)0.538
First, Do NOHARM (v2)
Stanford/Harvard consortium · 1,100 consultation cases- 1
Claude Opus 574.6%
- 2
Kimi K374.0%
- 3
GPT-5.6 Sol70.1%
- 4
GPT-5.570.0%
- 5
Gemini 3.1 Pro62.6%
HealthAgentBench
Microsoft Research · 54 agentic tasksCHI-Bench
actAVA · 75 operations workflowsMedCode (Vals AI)
Vals AI · 2,755 patient records- 1
Claude Opus 563.57%
- 2
Gemini 3.1 Pro Preview (02/26)59.06%
- 3
Claude Fable 556.07%
- 4
Gemini 3 Flash (12/25)55.92%
- 5
Gemini 3.5 Flash55.83%
MedScribe (Vals AI)
Vals AI · 100 SOAP-note cases- 1
Claude Opus 590.99%
- 2
Muse Spark 1.2~90
- 3
Muse Spark 1.1~90
- 4
Claude Fable 5~90
- 5
GPT 5.188.09%
MedXpertQA (MM)
Tsinghua University · 2,000 multimodal questions- 1
Gemini 3.1 Pro81.3%
- 2AQwen3.8 Max80.4%
- 3
Muse Spark78.4%
- 4
GPT-5.477.1%
- 5AQwen3.7 Plus71.0%
Artificial Analysis Healthcare & Medical Index
Artificial Analysis · 4-benchmark compositePhysicianBench
academic team · 100 clinical tasks- 1
GPT-5.546.3 ± 1.2
- 2
Claude Opus 4.631.7 ± 2.3
- 3
Claude Opus 4.729.3 ± 2.5
- 4
GPT-5.427.7 ± 1.5
- 5
Claude Sonnet 4.623.0 ± 2.6
EHR-Complex
academic team · 3,915-task test set- 1
GPT-5.4 (high reasoning)0.65
- 2
Gemini 3.1 Pro0.63
- 3
Kimi-K2.50.62
- 4AQwen3.5-397B0.62
- 5
GPT-5.4 (low reasoning)0.58
WHBench
academic team · 47 scenarios- 1
Claude Opus 4.672.1%
- 2
Claude Sonnet 4.667.1%
- 3
GPT-5.466.8%
- 4
Gemini 3 Flash Preview64.7%
- 5
GPT-4.151.8%
HealthAdminBench
Kinetic Systems · 135 admin tasks- 1
Claude Opus 4.6 (computer-use agent)36.3%
- 2
GPT-5.4 (computer-use agent)26.7%
- 3
Kimi K2.515.6%
- 4AQwen 3.513.3%
- 5
Gemini 3.1 Pro11.9%
OpenAI Dynamic Mental Health Evaluations
OpenAI · 3 safety metrics- 1
GPT-5.5 Instant (June Update)0.991
- 2
GPT-5.6 Sol (August)0.981
- 3
GPT-5.6 Luna (August)0.977
Scores appear exactly as each source publishes them. Two of the boards, HealthBench Professional and HealthBench Hard, are full leaderboards Arcophos runs.
How to read this index
Healthcare evaluation splintered as the frontier moved: rubric benchmarks, hard subsets, agentic clinical simulations, and coding or documentation tasks each measure something the others miss, and their scores live on incompatible scales. This index keeps them side by side without pretending they add up. For any benchmark, its page carries what it measures and where its numbers come from; for any model, its page collects every board it appears on. The sourcing rules are on the methodology page.
Models across boards
15 of the 70 tracked models hold results on more than one board. Each has a page gathering everything the index knows about it, listed on the models page.
What is not here
Some well-known names are absent on purpose. A benchmark leaves the index when frontier models stop being tested on it, and these are the notable cases with the reason each one dropped out:
| HealthBench Consensus | near-saturated physician-consensus baseline; frontier runs stopped reporting it separately |
|---|---|
| MedQA / MultiMedQA | exam-style multiple choice, saturated above 95 percent since 2025; archived by its trackers |
| AgentClinic | no public frontier-model results since 2025 |
| CRAFT-MD | no public frontier-model results since 2025 |
| MedAgentBench | v2 lives on inside the MAST composite; the standalone board has no current frontier rows |
| SDBench / MAI-DxO | Microsoft's 2025 sequential-diagnosis study was not re-run on current models |
| Open Medical-LLM Leaderboard (Hugging Face) | built on saturated exam sets; no frontier submissions in 2026 |
| MedArena | clinician preference arena; ratings pool too thin on current frontier models to quote |
| AMIE evaluations | Google DeepMind research prototypes, never opened to cross-vendor comparison |
| LiveClin, PrIME-LLM, MedMCP-Calc | single studies with two or fewer current-frontier rows; tracked for a future qualifying update |
Data
The index is downloadable in full as JSON and CSV, CC BY 4.0, from fixed paths. Citation guidance and version history live on the data page.