HealthAgentBench: current results
Microsoft Research · 54 tasks across 7 environments; 162 trials (3 attempts per task) · index updated August 16, 2026
Claude Code (Opus 5) holds the top current result on HealthAgentBench, 55% as of 2026-07, per HealthAgentBench leaderboard (Microsoft GitHub Pages). Agentic task success in realistic terminal-based healthcare environments built from real clinical artifacts; evaluates agent harnesses (Claude Code, Codex, Copilot) end to end, not bare models.
Current results
Result detail
| # | model | score | as of | |
|---|---|---|---|---|
| 1 | Claude Code (Opus 5) Anthropic $3.3/task; harness+model evaluated jointly | 55% | 2026-07 | |
| 2 | Codex (GPT-5.6-sol) OpenAI $5.2/task | 45% | 2026-07 | |
| 3 | Codex (GPT 5.5) OpenAI $2.8/task | 42% | 2026-07 | |
| 4 | Copilot (Opus 4.8) Microsoft/Anthropic $3.1/task | 36% | 2026-07 | |
| 5 | Copilot (GPT 5.5) Microsoft/OpenAI $2.6/task | 35% | 2026-07 | |
| 6 | Claude Code (Opus 4.8) Anthropic $4.0/task | 32% | 2026-07 | |
Scores appear exactly as HealthAgentBench leaderboard (Microsoft GitHub Pages) publishes them (independently run). Paper: arXiv 2606.31179. The rows are agent harnesses rather than bare models, and the board is run by Microsoft Research; Copilot, Microsoft's own harness, does not top it.
About the benchmark
| publisher | Microsoft Research |
|---|---|
| category | agentic and workflow benchmarks |
| released | 2026-07 |
| size | 54 tasks across 7 environments; 162 trials (3 attempts per task) |
| scale | mean task success rate, 0-100%, higher better; cost per task also reported |
| result basis | independently run |
| source | HealthAgentBench leaderboard (Microsoft GitHub Pages) |
| last frontier result | 2026-07 |
What is HealthAgentBench?
HealthAgentBench is a agentic and workflow benchmark from Microsoft Research, released 2026-07: 54 tasks across 7 environments; 162 trials (3 attempts per task), scored on a mean task success rate scale. Agentic task success in realistic terminal-based healthcare environments built from real clinical artifacts; evaluates agent harnesses (Claude Code, Codex, Copilot) end to end, not bare models.
Which model leads HealthAgentBench?
Claude Code (Opus 5) (Anthropic) holds the top current result on HealthAgentBench at 55%, per HealthAgentBench leaderboard (Microsoft GitHub Pages), as of 2026-07.
Where do the HealthAgentBench numbers come from?
From HealthAgentBench leaderboard (Microsoft GitHub Pages) (independently run). Paper: arXiv 2606.31179. The rows are agent harnesses rather than bare models, and the board is run by Microsoft Research; Copilot, Microsoft's own harness, does not top it.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.