PhysicianBench: current results
Academic team (Ruoqi Liu, Imran Q. Mohiuddin et al., arXiv 2605.02240); also a MAST component · 100 real-world clinical tasks, 21 specialties, 670 structured checkpoints (~27 tool calls per task) · index updated August 16, 2026
GPT-5.5 holds the top current result on PhysicianBench, 46.3 ± 1.2 as of 2026-05, per PhysicianBench paper (Table 2). LLM agents on long-horizon composite physician workflows inside real EHR environments, with execution-grounded verification against actual EHR systems via standard commercial APIs.
Current results
- 1
GPT-5.546.3 ± 1.2
- 2
Claude Opus 4.631.7 ± 2.3
- 3
Claude Opus 4.729.3 ± 2.5
- 4
GPT-5.427.7 ± 1.5
- 5
Claude Sonnet 4.623.0 ± 2.6
- 6
Kimi-K2.617.0 ± 2.6
- 7AQwen3.6-Plus13.7 ± 4.0
- 8
Gemini Pro 3.16.0 ± 1.0
Result detail
| # | model | score | as of | |
|---|---|---|---|---|
| 1 | GPT-5.5 OpenAI pass@1; Pass^3 28.0 | 46.3 ± 1.2 | 2026-05 | |
| 2 | Claude Opus 4.6 Anthropic Pass^3 18.0 | 31.7 ± 2.3 | 2026-05 | |
| 3 | Claude Opus 4.7 Anthropic Pass^3 18.0 | 29.3 ± 2.5 | 2026-05 | |
| 4 | GPT-5.4 OpenAI | 27.7 ± 1.5 | 2026-05 | |
| 5 | Claude Sonnet 4.6 Anthropic | 23.0 ± 2.6 | 2026-05 | |
| 6 | Kimi-K2.6 Moonshot AI open source | 17.0 ± 2.6 | 2026-05 | |
| 7 | A | Qwen3.6-Plus Alibaba | 13.7 ± 4.0 | 2026-05 |
| 8 | Gemini Pro 3.1 Google | 6.0 ± 1.0 | 2026-05 | |
Scores appear exactly as PhysicianBench paper (Table 2) publishes them (independently run). Scores come from the paper; there is no standalone public leaderboard, and the results also feed the ARISE MAST composite. The lower half of the table falls steeply, with several agents near 1 percent on Pass^3 consistency.
About the benchmark
| publisher | Academic team (Ruoqi Liu, Imran Q. Mohiuddin et al., arXiv 2605.02240); also a MAST component |
|---|---|
| category | agentic and workflow benchmarks |
| released | 2026-05 |
| size | 100 real-world clinical tasks, 21 specialties, 670 structured checkpoints (~27 tool calls per task) |
| scale | pass@1 success rate %, higher better (3 independent runs; Pass^3 also reported) |
| result basis | independently run |
| source | PhysicianBench paper (Table 2) |
| last frontier result | 2026-05 |
What is PhysicianBench?
PhysicianBench is a agentic and workflow benchmark from academic team, released 2026-05: 100 real-world clinical tasks, 21 specialties, 670 structured checkpoints (~27 tool calls per task), scored on a pass@1 success rate % scale. LLM agents on long-horizon composite physician workflows inside real EHR environments, with execution-grounded verification against actual EHR systems via standard commercial APIs.
Which model leads PhysicianBench?
GPT-5.5 (OpenAI) holds the top current result on PhysicianBench at 46.3 ± 1.2, per PhysicianBench paper (Table 2), as of 2026-05.
Where do the PhysicianBench numbers come from?
From PhysicianBench paper (Table 2) (independently run). Scores come from the paper; there is no standalone public leaderboard, and the results also feed the ARISE MAST composite. The lower half of the table falls steeply, with several agents near 1 percent on Pass^3 consistency.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.