Clinical Benchmarks

HealthAgentBench: current results

Microsoft Research · 54 tasks across 7 environments; 162 trials (3 attempts per task) · index updated August 16, 2026

Claude Code (Opus 5) holds the top current result on HealthAgentBench, 55% as of 2026-07, per HealthAgentBench leaderboard (Microsoft GitHub Pages). Agentic task success in realistic terminal-based healthcare environments built from real clinical artifacts; evaluates agent harnesses (Claude Code, Codex, Copilot) end to end, not bare models.

Current results

Result detail

#modelscoreas of
1Anthropic logoClaude Code (Opus 5) Anthropic
$3.3/task; harness+model evaluated jointly
55%2026-07
2OpenAI logoCodex (GPT-5.6-sol) OpenAI
$5.2/task
45%2026-07
3OpenAI logoCodex (GPT 5.5) OpenAI
$2.8/task
42%2026-07
4Microsoft/Anthropic logoCopilot (Opus 4.8) Microsoft/Anthropic
$3.1/task
36%2026-07
5Microsoft/OpenAI logoCopilot (GPT 5.5) Microsoft/OpenAI
$2.6/task
35%2026-07
6Anthropic logoClaude Code (Opus 4.8) Anthropic
$4.0/task
32%2026-07

Scores appear exactly as HealthAgentBench leaderboard (Microsoft GitHub Pages) publishes them (independently run). Paper: arXiv 2606.31179. The rows are agent harnesses rather than bare models, and the board is run by Microsoft Research; Copilot, Microsoft's own harness, does not top it.

About the benchmark

publisherMicrosoft Research
categoryagentic and workflow benchmarks
released2026-07
size54 tasks across 7 environments; 162 trials (3 attempts per task)
scalemean task success rate, 0-100%, higher better; cost per task also reported
result basisindependently run
sourceHealthAgentBench leaderboard (Microsoft GitHub Pages)
last frontier result2026-07

What is HealthAgentBench?

HealthAgentBench is a agentic and workflow benchmark from Microsoft Research, released 2026-07: 54 tasks across 7 environments; 162 trials (3 attempts per task), scored on a mean task success rate scale. Agentic task success in realistic terminal-based healthcare environments built from real clinical artifacts; evaluates agent harnesses (Claude Code, Codex, Copilot) end to end, not bare models.

Which model leads HealthAgentBench?

Claude Code (Opus 5) (Anthropic) holds the top current result on HealthAgentBench at 55%, per HealthAgentBench leaderboard (Microsoft GitHub Pages), as of 2026-07.

Where do the HealthAgentBench numbers come from?

From HealthAgentBench leaderboard (Microsoft GitHub Pages) (independently run). Paper: arXiv 2606.31179. The rows are agent harnesses rather than bare models, and the board is run by Microsoft Research; Copilot, Microsoft's own harness, does not top it.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.