Agentic and workflow benchmarks
5 tracked · updated August 16, 2026
Here models act instead of answering: multi-step work inside EHR systems, terminals, and administrative tools, scored on whether the task actually got done. Harness and model matter together, consistency across repeated runs is often the headline finding, and scores collapse fastest of any category when tasks get long.
HealthAgentBench
Microsoft Research · 54 agentic tasksCHI-Bench
actAVA · 75 operations workflowsPhysicianBench
academic team · 100 clinical tasks- 1
GPT-5.546.3 ± 1.2
- 2
Claude Opus 4.631.7 ± 2.3
- 3
Claude Opus 4.729.3 ± 2.5
- 4
GPT-5.427.7 ± 1.5
- 5
Claude Sonnet 4.623.0 ± 2.6
EHR-Complex
academic team · 3,915-task test set- 1
GPT-5.4 (high reasoning)0.65
- 2
Gemini 3.1 Pro0.63
- 3
Kimi-K2.50.62
- 4AQwen3.5-397B0.62
- 5
GPT-5.4 (low reasoning)0.58
HealthAdminBench
Kinetic Systems · 135 admin tasks- 1
Claude Opus 4.6 (computer-use agent)36.3%
- 2
GPT-5.4 (computer-use agent)26.7%
- 3
Kimi K2.515.6%
- 4AQwen 3.513.3%
- 5
Gemini 3.1 Pro11.9%
Which agentic and workflow benchmarks have current frontier-model results?
5 as of August 16, 2026: HealthAgentBench (Claude Code (Opus 5) leads at 55%); CHI-Bench (erius + claude-opus-5 leads at 54.7%); PhysicianBench (GPT-5.5 leads at 46.3 ± 1.2); EHR-Complex (GPT-5.4 (high reasoning) leads at 0.65); HealthAdminBench (Claude Opus 4.6 (computer-use agent) leads at 36.3%).
The other categories sit on the index: rubric-graded benchmarks, documentation and coding benchmarks, safety benchmarks, knowledge and exam benchmarks, composite indices.