Clinical Benchmarks

Healthcare AI Benchmarks: Current Frontier Results

Index updated August 16, 2026 · 17 benchmarks · 70 models

This index tracks every healthcare benchmark on which current frontier models hold a public result: 17 boards across 17 sources as of August 16, 2026, each score carrying the source it came from. No single model owns the field, though Claude Opus 5 leads more than one board. The table links to a page per benchmark and a page per model.

The benchmarks

One board per benchmark, bars scaled to that board's own leader. Scores sit on different scales and never compare across boards.

HealthBench Professional

OpenAI · 525 tasks
  1. 1Anthropic logoClaude Fable 50.660
  2. 2OpenAI logoGPT-5.6 Sol0.605
  3. 3Anthropic logoClaude Opus 50.598
via healthbenchprofessional.com · updated 2026-08full detail

HealthBench Hard

OpenAI · 1,000 conversations
  1. 1Meta logoMuse Spark0.428
  2. 2OpenAI logoGPT-5.6 Sol0.331
  3. 3OpenAI logoGPT-5.6 Terra0.327

HealthBench

OpenAI · 5,000 conversations
  1. 1Anthropic logoClaude Opus 567.1
  2. 2BBaichuan-M365.1
  3. 3OpenAI logoGPT-5.2-High63.3
  4. 4OpenAI logoGPT-5.6 Sol57.0
  5. 5OpenAI logoGPT-5.6 Terra57.0

MAST (Medical AI Superintelligence Test)

ARISE AI Research Network · 6-benchmark composite
  1. 1OpenAI logoGPT-5.6 Sol60.2%
  2. 2Moonshot AI logoKimi K360.1%
  3. 3Google logoGemini 3.6 Flash59.3%
  4. 4Google logoGemini 3.1 Pro58.9%
  5. 5AQwen3.5 397B A17B57.9%

MedHELM

Stanford CRFM · 121 tasks
  1. 1Google logoGemini 3.1 Pro (Preview)0.652
  2. 2Google logoGemini 3.5 Flash0.642
  3. 3Meta logoMuse Spark (2026-04-08)0.621
  4. 4OpenAI logoGPT-5.4 mini0.552
  5. 5OpenAI logoGPT-5.4 (2026-03-05)0.538

First, Do NOHARM (v2)

Stanford/Harvard consortium · 1,100 consultation cases
  1. 1Anthropic logoClaude Opus 574.6%
  2. 2Moonshot AI logoKimi K374.0%
  3. 3OpenAI logoGPT-5.6 Sol70.1%
  4. 4OpenAI logoGPT-5.570.0%
  5. 5Google logoGemini 3.1 Pro62.6%
via ARISE MAST technical leaderboard · updated 2026-08full detail

HealthAgentBench

Microsoft Research · 54 agentic tasks
  1. 1Anthropic logoClaude Code (Opus 5)55%
  2. 2OpenAI logoCodex (GPT-5.6-sol)45%
  3. 3OpenAI logoCodex (GPT 5.5)42%
  4. 4Microsoft/Anthropic logoCopilot (Opus 4.8)36%
  5. 5Microsoft/OpenAI logoCopilot (GPT 5.5)35%

CHI-Bench

actAVA · 75 operations workflows
  1. 1Humana (harness) / Anthropic (model) logoerius + claude-opus-554.7%
  2. 2Humana / Anthropic logoerius + claude-opus-4-837.3%
  3. 3Anthropic logoclaude-code + claude-opus-537.3%
  4. 4Anthropic logoclaude-code + claude-opus-4-833.3%
  5. 5Anthropic logoclaude-code + claude-opus-4-628.0%

MedCode (Vals AI)

Vals AI · 2,755 patient records
  1. 1Anthropic logoClaude Opus 563.57%
  2. 2Google logoGemini 3.1 Pro Preview (02/26)59.06%
  3. 3Anthropic logoClaude Fable 556.07%
  4. 4Google logoGemini 3 Flash (12/25)55.92%
  5. 5Google logoGemini 3.5 Flash55.83%
via Vals AI MedCode leaderboard · updated 2026-08full detail

MedScribe (Vals AI)

Vals AI · 100 SOAP-note cases
  1. 1Anthropic logoClaude Opus 590.99%
  2. 2Meta logoMuse Spark 1.2~90
  3. 3Meta logoMuse Spark 1.1~90
  4. 4Anthropic logoClaude Fable 5~90
  5. 5OpenAI logoGPT 5.188.09%
via Vals AI MedScribe leaderboard · updated 2026-08all 6 results

MedXpertQA (MM)

Tsinghua University · 2,000 multimodal questions
  1. 1Google logoGemini 3.1 Pro81.3%
  2. 2AQwen3.8 Max80.4%
  3. 3Meta logoMuse Spark78.4%
  4. 4OpenAI logoGPT-5.477.1%
  5. 5AQwen3.7 Plus71.0%

Artificial Analysis Healthcare & Medical Index

Artificial Analysis · 4-benchmark composite
  1. 1Anthropic logoClaude Opus 5 (Adaptive Reasoning, Max Effort)51
  2. 2Anthropic logoClaude Opus 5 (Adaptive Reasoning, Xhigh Effort)51
  3. 3Anthropic logoClaude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)51

PhysicianBench

academic team · 100 clinical tasks
  1. 1OpenAI logoGPT-5.546.3 ± 1.2
  2. 2Anthropic logoClaude Opus 4.631.7 ± 2.3
  3. 3Anthropic logoClaude Opus 4.729.3 ± 2.5
  4. 4OpenAI logoGPT-5.427.7 ± 1.5
  5. 5Anthropic logoClaude Sonnet 4.623.0 ± 2.6

EHR-Complex

academic team · 3,915-task test set
  1. 1OpenAI logoGPT-5.4 (high reasoning)0.65
  2. 2Google logoGemini 3.1 Pro0.63
  3. 3Moonshot AI logoKimi-K2.50.62
  4. 4AQwen3.5-397B0.62
  5. 5OpenAI logoGPT-5.4 (low reasoning)0.58

WHBench

academic team · 47 scenarios
  1. 1Anthropic logoClaude Opus 4.672.1%
  2. 2Anthropic logoClaude Sonnet 4.667.1%
  3. 3OpenAI logoGPT-5.466.8%
  4. 4Google logoGemini 3 Flash Preview64.7%
  5. 5OpenAI logoGPT-4.151.8%

HealthAdminBench

Kinetic Systems · 135 admin tasks
  1. 1Anthropic logoClaude Opus 4.6 (computer-use agent)36.3%
  2. 2OpenAI logoGPT-5.4 (computer-use agent)26.7%
  3. 3Moonshot AI logoKimi K2.515.6%
  4. 4AQwen 3.513.3%
  5. 5Google logoGemini 3.1 Pro11.9%

OpenAI Dynamic Mental Health Evaluations

OpenAI · 3 safety metrics
  1. 1OpenAI logoGPT-5.5 Instant (June Update)0.991
  2. 2OpenAI logoGPT-5.6 Sol (August)0.981
  3. 3OpenAI logoGPT-5.6 Luna (August)0.977
via OpenAI August Updates · updated 2026-08full detail

Scores appear exactly as each source publishes them. Two of the boards, HealthBench Professional and HealthBench Hard, are full leaderboards Arcophos runs.

17
benchmarks tracked
70
frontier models covered
13
labs represented
17
distinct sources
2
boards run by Arcophos
15
models on multiple boards

How to read this index

Healthcare evaluation splintered as the frontier moved: rubric benchmarks, hard subsets, agentic clinical simulations, and coding or documentation tasks each measure something the others miss, and their scores live on incompatible scales. This index keeps them side by side without pretending they add up. For any benchmark, its page carries what it measures and where its numbers come from; for any model, its page collects every board it appears on. The sourcing rules are on the methodology page.

Models across boards

15 of the 70 tracked models hold results on more than one board. Each has a page gathering everything the index knows about it, listed on the models page.

What is not here

Some well-known names are absent on purpose. A benchmark leaves the index when frontier models stop being tested on it, and these are the notable cases with the reason each one dropped out:

HealthBench Consensusnear-saturated physician-consensus baseline; frontier runs stopped reporting it separately
MedQA / MultiMedQAexam-style multiple choice, saturated above 95 percent since 2025; archived by its trackers
AgentClinicno public frontier-model results since 2025
CRAFT-MDno public frontier-model results since 2025
MedAgentBenchv2 lives on inside the MAST composite; the standalone board has no current frontier rows
SDBench / MAI-DxOMicrosoft's 2025 sequential-diagnosis study was not re-run on current models
Open Medical-LLM Leaderboard (Hugging Face)built on saturated exam sets; no frontier submissions in 2026
MedArenaclinician preference arena; ratings pool too thin on current frontier models to quote
AMIE evaluationsGoogle DeepMind research prototypes, never opened to cross-vendor comparison
LiveClin, PrIME-LLM, MedMCP-Calcsingle studies with two or fewer current-frontier rows; tracked for a future qualifying update

Data

The index is downloadable in full as JSON and CSV, CC BY 4.0, from fixed paths. Citation guidance and version history live on the data page.