Claude Opus 4.6: healthcare benchmark results
Anthropic · 5 boards · updated August 16, 2026
The index currently holds 5 results for Claude Opus 4.6: 0.456 on MedHELM (8 of 11), 86.74% on MedScribe (Vals AI) (6 of 84), 64.8% on MedXpertQA (MM) (7 of 8), 31.7 ± 2.3 on PhysicianBench (2 of 13), 72.1% on WHBench (1 of 22). It tops WHBench. Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.
Results by benchmark
| benchmark | score | position | as of |
|---|---|---|---|
| MedHELM via MedHELM leaderboard (medhelm.org), v5.0.0 | 0.456 | 8 of 11 | 2026-05 |
| MedScribe (Vals AI) via Vals AI MedScribe leaderboard | 86.74% | 6 of 84 | 2026-02 |
| MedXpertQA (MM) via benchlm.ai mirror of Meta's Muse Spark evaluation | 64.8% | 7 of 8 | 2026-08 |
| PhysicianBench via PhysicianBench paper (Table 2) | 31.7 ± 2.3 | 2 of 13 | 2026-05 |
| WHBench via arXiv paper (v2 revised 2026-07-23) | 72.1% | 1 of 22 | 2026-07 |
Position counts against the source's full board, including rows this index does not mirror. Config caveats, where a source noted any, are on each benchmark's page. Sources also list this model as "Claude 4.6 Opus".
Which healthcare benchmarks is Claude Opus 4.6 scored on?
As of August 16, 2026, Claude Opus 4.6 holds current results on 5 tracked benchmarks: MedHELM, MedScribe (Vals AI), MedXpertQA (MM), PhysicianBench, WHBench.
How does Claude Opus 4.6 rank on them?
Claude Opus 4.6 stands at 0.456 on MedHELM (8 of 11), 86.74% on MedScribe (Vals AI) (6 of 84), 64.8% on MedXpertQA (MM) (7 of 8), 31.7 ± 2.3 on PhysicianBench (2 of 13), 72.1% on WHBench (1 of 22). It holds first place on WHBench.
The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.