# Clinical Benchmarks index, 2026-08-16 18 healthcare AI benchmarks currently carry public results for frontier models. Each section below gives one benchmark's current results exactly as its named source publishes them. Scores are never comparable across benchmarks. Source: https://clinicalbenchmarks.ai ## HealthBench Professional 525 tasks that physicians picked out of 15,079 real workplace AI conversations, spanning care consults, clinical documentation, and medical research, each judged on a rubric physicians wrote for it. Published by OpenAI (2026-04); 525 physician-authored tasks; scale 0 to 1, higher is better. Source: healthbenchprofessional.com (independent-run), https://healthbenchprofessional.com. Full leaderboard: https://healthbenchprofessional.com | # | model | lab | score | as of | |---|---|---|---|---| | 1 | Claude Fable 5 | Anthropic | 0.660 | 2026-08 | | 2 | GPT-5.6 Sol | OpenAI | 0.605 | 2026-08 | | 3 | Claude Opus 5 | Anthropic | 0.598 | 2026-08 | Grader GPT-5.4 at low reasoning effort with a length adjustment. Physician-written responses score 0.437 on the same rubrics. ## HealthBench Hard The bottom fifth of HealthBench: 1,000 conversations where frontier models failed most at the May 2025 release, still graded on the original physician-written rubrics. Published by OpenAI (2025-05); 1,000 conversations; scale 0 to 1, higher is better. Source: healthbenchhard.ai (mixed), https://healthbenchhard.ai. Full leaderboard: https://healthbenchhard.ai | # | model | lab | score | as of | |---|---|---|---|---| | 1 | Muse Spark | Meta | 0.428 | 2026-08 | | 2 | GPT-5.6 Sol | OpenAI | 0.331 | 2026-08 | | 3 | GPT-5.6 Terra | OpenAI | 0.327 | 2026-08 | Default grader GPT-4.1. Best score at the May 2025 release was o3's 0.320. ## HealthBench 5,000 realistic multi-turn health conversations graded against physician-written rubrics (48,562 criteria) covering accuracy, completeness, context awareness, communication, and instruction following. OpenAI now also reports a length-adjusted variant that penalizes verbosity. Published by OpenAI (2025-05); 5,000 conversations; scale 0-100 rubric-point percentage (some sites display 0-1), higher better; length-adjusted and unadjusted variants. Source: OpenAI Deployment Safety Hub (GPT-5.6 system card + August 2026 updates); benchlm.ai and llm-stats.com mirror (mixed), https://deploymentsafety.openai.com/gpt-5-6-preview/healthbench | # | model | lab | score | as of | |---|---|---|---|---| | 1 | Claude Opus 5 | Anthropic | 67.1 | 2026-08 | | 2 | Baichuan-M3 | Baichuan | 65.1 | 2026-02 | | 3 | GPT-5.2-High | OpenAI | 63.3 | 2026-02 | | 4 | GPT-5.6 Sol | OpenAI | 57.0 | 2026-06 | | 5 | GPT-5.6 Terra | OpenAI | 57.0 | 2026-06 | | 6 | GPT-5.5 | OpenAI | 56.5 | 2026-06 | | 7 | GPT-5.6 Luna | OpenAI | 55.8 | 2026-06 | | 8 | GPT-5.6 Sol (August) | OpenAI | 55.0 | 2026-08 | Cross-source rows are not directly comparable: Anthropic reports a raw score under its own protocol, OpenAI leads with length-adjusted numbers at maximum reasoning effort and reports production Instant settings separately, and Baichuan ran its competitors itself. The benchmark's official grader is GPT-4.1. OpenAI has said the parent set is approaching a noise ceiling for frontier models and points to HealthBench Professional for continued measurement. ## Health Optimization Bench Frontier models on hard, freshness-dependent questions in preventive and optimization medicine. Every task is written against a primary source, audited by model families that did not author it, and scored blind by a panel of independent families. The v1 release set covers incretin therapeutics evidence. Published by Health Optimization Bench (2026-08); 89 released tasks (v1 evidence suite); scale 0-100 rubric credit, higher better. Source: healthoptimizationbench.com (independent-run), https://healthoptimizationbench.com. Full leaderboard: https://healthoptimizationbench.com | # | model | lab | score | as of | |---|---|---|---|---| | 1 | Claude Fable 5 | Anthropic | 83.8 | 2026-08 | | 2 | Grok 4.6 | xAI | 81.3 | 2026-08 | | 3 | Claude Opus 5 | Anthropic | 78.3 | 2026-08 | Cross-family authoring with blind three-family panel grading; the authoring family never grades its own task, and 95 percent bootstrap confidence intervals accompany every score on the site. ## MAST (Medical AI Superintelligence Test) Composite score across curated clinical benchmarks spanning diagnostic reasoning, management reasoning, safety, multimodal images, multimodal radiology, and agentic capability. Components: First Do NOHARM v2, SCT-Bench, MedAgentBench v2, PhysicianBench, ReXrank Mini, CPC-Bench. Published by ARISE AI Research Network (multi-institutional) (2026-08); composite of 6 component benchmarks; 11 models; scale percentage composite, higher better. Source: ARISE MAST leaderboard (independent-run), https://arise-ai.org/mast | # | model | lab | score | as of | |---|---|---|---|---| | 1 | GPT-5.6 Sol | OpenAI | 60.2% | 2026-08 | | 2 | Kimi K3 | Moonshot AI | 60.1% | 2026-08 | | 3 | Gemini 3.6 Flash | Google | 59.3% | 2026-08 | | 4 | Gemini 3.1 Pro | Google | 58.9% | 2026-08 | | 5 | Qwen3.5 397B A17B | Alibaba | 57.9% | 2026-08 | | 6 | Claude Opus 5 | Anthropic | 57.1% | 2026-08 | | 7 | Claude Sonnet 5 | Anthropic | 56.6% | 2026-08 | | 8 | Grok 4.3 | xAI | 53.7% | 2026-08 | The board is marked as a preview and was last updated August 15, 2026; component-level breakdowns are published only for First, Do NOHARM v2. Scores may move before the full release. ## MedHELM Holistic evaluation of LLMs on 121 clinical tasks across 5 categories and 22 subcategories (31 datasets) in a clinician-validated taxonomy; ranked by mean win rate. Published by Stanford CRFM / HAI and multi-institution collaborators (2025-02); 121 tasks / 31 datasets; scale mean win rate 0-1, higher better. Source: MedHELM leaderboard (medhelm.org), v5.0.0 (official-leaderboard), https://medhelm.org/ | # | model | lab | score | as of | |---|---|---|---|---| | 1 | Gemini 3.1 Pro (Preview) | Google | 0.652 | 2026-05 | | 2 | Gemini 3.5 Flash | Google | 0.642 | 2026-05 | | 3 | Muse Spark (2026-04-08) | Meta | 0.621 | 2026-05 | | 4 | GPT-5.4 mini | OpenAI | 0.552 | 2026-05 | | 5 | GPT-5.4 (2026-03-05) | OpenAI | 0.538 | 2026-05 | | 6 | Gemini 2.5 Pro | Google | 0.529 | 2026-05 | | 7 | DeepSeek R1 | DeepSeek | 0.485 | 2026-05 | | 8 | Claude 4.6 Opus | Anthropic | 0.456 | 2026-05 | | 9 | Claude 3.7 Sonnet | Anthropic | 0.45 | 2026-05 | | 10 | Gemini 2.0 Flash | Google | 0.342 | 2026-05 | Version 5.0.0, last updated May 14, 2026, run by the Stanford-led maintainers on a roughly quarterly cadence. No Claude 5 family or GPT-5.6 rows yet. Mean win rate is relative to the evaluated cohort, so scores shift whenever the model set changes. ## First, Do NOHARM (v2) Frequency and severity of potentially harmful errors in LLM-generated medical consultation recommendations (Numerous Options Harm Assessment for Risk in Medicine); primary-care-to-specialist consults. Published by Stanford/Harvard-led consortium (50+ researchers incl. 29 board-certified physicians); hosted by ARISE (2025-12); 1,100 consultation cases, 10 specialties, 12,747 expert annotations on 4,249 management options; scale percentage safety score, higher better. Source: ARISE MAST technical leaderboard (official-leaderboard), https://arise-ai.org/mast/technical | # | model | lab | score | as of | |---|---|---|---|---| | 1 | Claude Opus 5 | Anthropic | 74.6% | 2026-08 | | 2 | Kimi K3 | Moonshot AI | 74.0% | 2026-08 | | 3 | GPT-5.6 Sol | OpenAI | 70.1% | 2026-08 | | 4 | GPT-5.5 | OpenAI | 70.0% | 2026-08 | | 5 | Gemini 3.1 Pro | Google | 62.6% | 2026-08 | Paper: arXiv 2512.01241. The v1 study found potential for severe harm in up to 24.6 percent of directly applied recommendations, with errors of omission behind more than 80 percent of the severe cases. ## HealthAgentBench Agentic task success in realistic terminal-based healthcare environments built from real clinical artifacts; evaluates agent harnesses (Claude Code, Codex, Copilot) end to end, not bare models. Published by Microsoft Research (2026-07); 54 tasks across 7 environments; 162 trials (3 attempts per task); scale mean task success rate, 0-100%, higher better; cost per task also reported. Source: HealthAgentBench leaderboard (Microsoft GitHub Pages) (independent-run), https://microsoft.github.io/HealthAgentBench/ | # | model | lab | score | as of | |---|---|---|---|---| | 1 | Claude Code (Opus 5) | Anthropic | 55% | 2026-07 | | 2 | Codex (GPT-5.6-sol) | OpenAI | 45% | 2026-07 | | 3 | Codex (GPT 5.5) | OpenAI | 42% | 2026-07 | | 4 | Copilot (Opus 4.8) | Microsoft/Anthropic | 36% | 2026-07 | | 5 | Copilot (GPT 5.5) | Microsoft/OpenAI | 35% | 2026-07 | | 6 | Claude Code (Opus 4.8) | Anthropic | 32% | 2026-07 | Paper: arXiv 2606.31179. The rows are agent harnesses rather than bare models, and the board is run by Microsoft Research; Copilot, Microsoft's own harness, does not top it. ## CHI-Bench Long-horizon US healthcare operations workflows for agents (prior authorization, utilization management, care management), 60-80 step tasks across 4-6 stages, judged by deterministic unit tests plus an LLM judge for evidence grounding, consent, and cross-stage consistency. Published by actAVA.ai (2026-05); 75 workflows (25 per domain), 21 healthcare applications, 200+ MCP tools; scale pass@1 with binary 0/1 reward, higher better. Source: CHI-Bench leaderboard (actAVA) (mixed), https://actava.ai/benchmarks/leaderboards | # | model | lab | score | as of | |---|---|---|---|---| | 1 | erius + claude-opus-5 | Humana (harness) / Anthropic (model) | 54.7% | 2026-08 | | 2 | erius + claude-opus-4-8 | Humana / Anthropic | 37.3% | 2026-08 | | 3 | claude-code + claude-opus-5 | Anthropic | 37.3% | 2026-08 | | 4 | claude-code + claude-opus-4-8 | Anthropic | 33.3% | 2026-08 | | 5 | claude-code + claude-opus-4-6 | Anthropic | 28.0% | 2026-08 | | 6 | claude-code + claude-sonnet-4-6 | Anthropic | 26.2% | 2026-08 | | 7 | codex + gpt-5.6-sol | OpenAI | 25.3% | 2026-08 | | 8 | openai-agents + kimi-k3 | Moonshot AI | 25.3% | 2026-08 | | 9 | claude-code + claude-opus-4-7 | Anthropic | 24.4% | 2026-08 | | 10 | claude-code + claude-fable-5 | Anthropic | 24.0% | 2026-08 | Released May 20, 2026; updated August 12, 2026 across 45 harness configurations. The launch report led with reliability, not capability: no agent stayed above 20 percent across three identical runs. Harness choice matters as much as model choice, and the board accepts community submissions, so rows mix author-run and submitted results. ## MedCode (Vals AI) ICD-10-CM diagnosis coding for entire hospital stays: models assign primary and secondary codes from discharge summaries plus progress/consult notes; ground truth double-annotated by certified professional coders. Published by Vals AI (dataset with Protege) (2026-02); 2,755 patient records; scale percentage accuracy 0-100, higher better. Source: Vals AI MedCode leaderboard (independent-run), https://www.vals.ai/benchmarks/medcode | # | model | lab | score | as of | |---|---|---|---|---| | 1 | Claude Opus 5 | Anthropic | 63.57% | 2026-08 | | 2 | Gemini 3.1 Pro Preview (02/26) | Google | 59.06% | 2026-08 | | 3 | Claude Fable 5 | Anthropic | 56.07% | 2026-08 | | 4 | Gemini 3 Flash (12/25) | Google | 55.92% | 2026-08 | | 5 | Gemini 3.5 Flash | Google | 55.83% | 2026-08 | Vals AI runs every model itself; 85 models on the board, last updated August 15, 2026. Vals AI pairs this board with MedScribe, and its writeup notes that coding accuracy lags documentation quality. ## MedScribe (Vals AI) Clinical documentation support: quality of SOAP notes generated from clinical visits, scored against rubrics for documentation quality and compliance. Published by Vals AI (dataset with Protege) (2026-02); 100 rubric-scored SOAP-note cases; scale percentage accuracy 0-100, higher better. Source: Vals AI MedScribe leaderboard (independent-run), https://www.vals.ai/benchmarks/medscribe | # | model | lab | score | as of | |---|---|---|---|---| | 1 | Claude Opus 5 | Anthropic | 90.99% | 2026-08 | | 2 | Muse Spark 1.2 | Meta | ~90 | 2026-08 | | 3 | Muse Spark 1.1 | Meta | ~90 | 2026-08 | | 4 | Claude Fable 5 | Anthropic | ~90 | 2026-08 | | 5 | GPT 5.1 | OpenAI | 88.09% | 2026-02 | | 6 | Claude Opus 4.6 | Anthropic | 86.74% | 2026-02 | Vals AI self-runs; 84 models, last updated August 15, 2026. Top scores cluster near 90, so the leaders sit close to the ceiling, and exact decimals below first place are not always displayed. ## MedXpertQA (MM) Expert-level multimodal medical multiple-choice QA covering clinical images (X-ray, histology, dermatology, charts) across 17 specialties; MM subset of the 4,460-question MedXpertQA benchmark. Published by TsinghuaC3I (Tsinghua University) (2025-01); 2,000 multimodal questions (MM subset); scale percentage accuracy 0-100, higher better. Source: benchlm.ai mirror of Meta's Muse Spark evaluation (vendor-reported), https://benchlm.ai/benchmarks/medxpertqamm | # | model | lab | score | as of | |---|---|---|---|---| | 1 | Gemini 3.1 Pro | Google | 81.3% | 2026-08 | | 2 | Qwen3.8 Max | Alibaba | 80.4% | 2026-08 | | 3 | Muse Spark | Meta | 78.4% | 2026-08 | | 4 | GPT-5.4 | OpenAI | 77.1% | 2026-08 | | 5 | Qwen3.7 Plus | Alibaba | 71.0% | 2026-08 | | 6 | Grok 4.20 | xAI | 65.8% | 2026-08 | | 7 | Claude Opus 4.6 | Anthropic | 64.8% | 2026-08 | | 8 | Gemma 4 12B | Google | 48.7% | 2026-08 | The current frontier table is vendor-reported, from Meta's Muse Spark launch evaluation, mirrored display-only on benchlm.ai. The Text subset has no comparable current table. ## Artificial Analysis Healthcare & Medical Index Weighted composite for healthcare and medical work: Medical & Health Knowledge 35%, Agentic Knowledge Work 25%, Non-Hallucination 15%, Reasoning 15%, Agentic Customer Interaction 10%, drawn from AA-Omniscience, GDPval-AA v2, Humanity's Last Exam, and tau3-Banking. Published by Artificial Analysis (2026-08); 27 of 159 models scored; composite of 4 underlying benchmarks; scale index score, higher better. Source: Artificial Analysis healthcare capability page (independent-run), https://artificialanalysis.ai/models/capabilities/healthcare-and-medical | # | model | lab | score | as of | |---|---|---|---|---| | 1 | Claude Opus 5 (Adaptive Reasoning, Max Effort) | Anthropic | 51 | 2026-08 | | 2 | Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) | Anthropic | 51 | 2026-08 | | 3 | Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) | Anthropic | 51 | 2026-08 | A weighted composite of general-purpose benchmarks tilted toward health-relevant slices rather than purpose-built clinical tasks, which is why this index files it as a capability index. The page differentiates only its top scores numerically. ## PhysicianBench LLM agents on long-horizon composite physician workflows inside real EHR environments, with execution-grounded verification against actual EHR systems via standard commercial APIs. Published by Academic team (Ruoqi Liu, Imran Q. Mohiuddin et al., arXiv 2605.02240); also a MAST component (2026-05); 100 real-world clinical tasks, 21 specialties, 670 structured checkpoints (~27 tool calls per task); scale pass@1 success rate %, higher better (3 independent runs; Pass^3 also reported). Source: PhysicianBench paper (Table 2) (independent-run), https://arxiv.org/abs/2605.02240 | # | model | lab | score | as of | |---|---|---|---|---| | 1 | GPT-5.5 | OpenAI | 46.3 ± 1.2 | 2026-05 | | 2 | Claude Opus 4.6 | Anthropic | 31.7 ± 2.3 | 2026-05 | | 3 | Claude Opus 4.7 | Anthropic | 29.3 ± 2.5 | 2026-05 | | 4 | GPT-5.4 | OpenAI | 27.7 ± 1.5 | 2026-05 | | 5 | Claude Sonnet 4.6 | Anthropic | 23.0 ± 2.6 | 2026-05 | | 6 | Kimi-K2.6 | Moonshot AI | 17.0 ± 2.6 | 2026-05 | | 7 | Qwen3.6-Plus | Alibaba | 13.7 ± 4.0 | 2026-05 | | 8 | Gemini Pro 3.1 | Google | 6.0 ± 1.0 | 2026-05 | Scores come from the paper; there is no standalone public leaderboard, and the results also feed the ARISE MAST composite. The lower half of the table falls steeply, with several agents near 1 percent on Pass^3 consistency. ## EHR-Complex Agentic clinical reasoning over MIMIC-IV EHR databases via SQL and Python across six clinical intents, at patient and population level with temporal evidence paths. Published by Academic team (Qiao et al., Ant Group-affiliated; arXiv 2606.23301) (2026-06); ~52,000 tasks (3,915-task test set) over 365K patients, 31 tables, 500M+ records; scale exact-match accuracy, 0-1, higher better. Source: EHR-Complex paper (independent-run), https://arxiv.org/abs/2606.23301 | # | model | lab | score | as of | |---|---|---|---|---| | 1 | GPT-5.4 (high reasoning) | OpenAI | 0.65 | 2026-06 | | 2 | Gemini 3.1 Pro | Google | 0.63 | 2026-06 | | 3 | Kimi-K2.5 | Moonshot AI | 0.62 | 2026-06 | | 4 | Qwen3.5-397B | Alibaba | 0.62 | 2026-06 | | 5 | GPT-5.4 (low reasoning) | OpenAI | 0.58 | 2026-06 | | 6 | Claude Sonnet 4.6 | Anthropic | 0.36 | 2026-06 | Scores come from the paper, which reports both a headline 12-model evaluation and human-validated configurations; rows here mix the two, labeled in the config column. Consistency drops below 50 percent at Pass^4 for nearly every model. ## WHBench Women's health: 47 expert-crafted scenarios across 10 topics graded on a 23-criterion rubric for clinical accuracy, safety, equity, and guideline adherence; targets failure modes like outdated guidelines, unsafe omissions, dosing errors, equity blind spots. Published by Independent researchers (Maurya, Govindgari, Kumar) (2026-04); 47 scenarios / 3,100 scored responses across 22 models; scale mean normalized percentage 0-100, higher better. Source: arXiv paper (v2 revised 2026-07-23) (independent-run), https://arxiv.org/abs/2604.00024 | # | model | lab | score | as of | |---|---|---|---|---| | 1 | Claude Opus 4.6 | Anthropic | 72.1% | 2026-07 | | 2 | Claude Sonnet 4.6 | Anthropic | 67.1% | 2026-07 | | 3 | GPT-5.4 | OpenAI | 66.8% | 2026-07 | | 4 | Gemini 3 Flash Preview | Google | 64.7% | 2026-07 | | 5 | GPT-4.1 | OpenAI | 51.8% | 2026-07 | | 6 | GPT-4o | OpenAI | 44.6% | 2026-07 | An academic study with expert validation rather than a live leaderboard; the model set was frozen in March 2026, before GPT-5.6 and the Claude 5 family shipped. ## HealthAdminBench End-to-end task success of computer-use LLM agents on healthcare administration workflows: prior authorizations, denial appeals, and DME ordering; success requires completing every subtask in a task. Published by Kinetic Systems (with Stanford Hospital domain experts) (2026-04); 135 tasks / 1,698 rubric-scored subtasks; scale percentage end-to-end task success 0-100, higher better. Source: Kinetic Systems blog (independent-run), https://kineticsystems.ai/blog/healthadminbench-automating-healthcare-administration-with-computer-use-agents | # | model | lab | score | as of | |---|---|---|---|---| | 1 | Claude Opus 4.6 (computer-use agent) | Anthropic | 36.3% | 2026-04 | | 2 | GPT-5.4 (computer-use agent) | OpenAI | 26.7% | 2026-04 | | 3 | Kimi K2.5 | Moonshot AI | 15.6% | 2026-04 | | 4 | Qwen 3.5 | Alibaba | 13.3% | 2026-04 | | 5 | Gemini 3.1 Pro | Google | 11.9% | 2026-04 | Run by Kinetic Systems' research team: independent of the model vendors, though published by a company selling healthcare-admin automation. No GPT-5.6 or Claude 5 rows yet, and no stated refresh cadence. ## OpenAI Dynamic Mental Health Evaluations Multi-turn adversarial user simulations for mental health, emotional reliance, and self-harm response quality, where conversations evolve in response to model outputs rather than following fixed scripts. Published by OpenAI (2026-08); dynamic simulated conversations (counts not disclosed); scale compliance rate per metric, 0 to 1, higher better; headline number is the mental-health metric. Source: OpenAI GPT-5.6 August Updates (PDF) (vendor-reported), https://cdn.openai.com/pdf/GPT_5_6_August_Updates.pdf | # | model | lab | score | as of | |---|---|---|---|---| | 1 | GPT-5.5 Instant (June Update) | OpenAI | 0.991 | 2026-08 | | 2 | GPT-5.6 Sol (August) | OpenAI | 0.981 | 2026-08 | | 3 | GPT-5.6 Luna (August) | OpenAI | 0.977 | 2026-08 | An internal OpenAI safety evaluation covering OpenAI models only; it is not independently runnable, and OpenAI notes the error rates are not representative of average production traffic. Listed as a vendor safety eval, not a cross-vendor benchmark. ## Reuse Index data: CC BY 4.0. Cite the index as: Clinical Benchmarks index, 2026-08-16 snapshot. https://clinicalbenchmarks.ai. Cite individual scores to the source named in their section. Machine-readable: https://clinicalbenchmarks.ai/data/benchmarks.json