Clinical Benchmarks

PhysicianBench: current results

Academic team (Ruoqi Liu, Imran Q. Mohiuddin et al., arXiv 2605.02240); also a MAST component · 100 real-world clinical tasks, 21 specialties, 670 structured checkpoints (~27 tool calls per task) · index updated August 16, 2026

GPT-5.5 holds the top current result on PhysicianBench, 46.3 ± 1.2 as of 2026-05, per PhysicianBench paper (Table 2). LLM agents on long-horizon composite physician workflows inside real EHR environments, with execution-grounded verification against actual EHR systems via standard commercial APIs.

Current results

  1. 1OpenAI logoGPT-5.546.3 ± 1.2
  2. 2Anthropic logoClaude Opus 4.631.7 ± 2.3
  3. 3Anthropic logoClaude Opus 4.729.3 ± 2.5
  4. 4OpenAI logoGPT-5.427.7 ± 1.5
  5. 5Anthropic logoClaude Sonnet 4.623.0 ± 2.6
  6. 6Moonshot AI logoKimi-K2.617.0 ± 2.6
  7. 7AQwen3.6-Plus13.7 ± 4.0
  8. 8Google logoGemini Pro 3.16.0 ± 1.0

Result detail

#modelscoreas of
1OpenAI logoGPT-5.5 OpenAI
pass@1; Pass^3 28.0
46.3 ± 1.22026-05
2Anthropic logoClaude Opus 4.6 Anthropic
Pass^3 18.0
31.7 ± 2.32026-05
3Anthropic logoClaude Opus 4.7 Anthropic
Pass^3 18.0
29.3 ± 2.52026-05
4OpenAI logoGPT-5.4 OpenAI27.7 ± 1.52026-05
5Anthropic logoClaude Sonnet 4.6 Anthropic23.0 ± 2.62026-05
6Moonshot AI logoKimi-K2.6 Moonshot AI
open source
17.0 ± 2.62026-05
7AQwen3.6-Plus Alibaba13.7 ± 4.02026-05
8Google logoGemini Pro 3.1 Google6.0 ± 1.02026-05

Scores appear exactly as PhysicianBench paper (Table 2) publishes them (independently run). Scores come from the paper; there is no standalone public leaderboard, and the results also feed the ARISE MAST composite. The lower half of the table falls steeply, with several agents near 1 percent on Pass^3 consistency.

About the benchmark

publisherAcademic team (Ruoqi Liu, Imran Q. Mohiuddin et al., arXiv 2605.02240); also a MAST component
categoryagentic and workflow benchmarks
released2026-05
size100 real-world clinical tasks, 21 specialties, 670 structured checkpoints (~27 tool calls per task)
scalepass@1 success rate %, higher better (3 independent runs; Pass^3 also reported)
result basisindependently run
sourcePhysicianBench paper (Table 2)
last frontier result2026-05

What is PhysicianBench?

PhysicianBench is a agentic and workflow benchmark from academic team, released 2026-05: 100 real-world clinical tasks, 21 specialties, 670 structured checkpoints (~27 tool calls per task), scored on a pass@1 success rate % scale. LLM agents on long-horizon composite physician workflows inside real EHR environments, with execution-grounded verification against actual EHR systems via standard commercial APIs.

Which model leads PhysicianBench?

GPT-5.5 (OpenAI) holds the top current result on PhysicianBench at 46.3 ± 1.2, per PhysicianBench paper (Table 2), as of 2026-05.

Where do the PhysicianBench numbers come from?

From PhysicianBench paper (Table 2) (independently run). Scores come from the paper; there is no standalone public leaderboard, and the results also feed the ARISE MAST composite. The lower half of the table falls steeply, with several agents near 1 percent on Pass^3 consistency.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.