HealthAdminBench: current results
Kinetic Systems (with Stanford Hospital domain experts) · 135 tasks / 1,698 rubric-scored subtasks · index updated August 16, 2026
Claude Opus 4.6 (computer-use agent) holds the top current result on HealthAdminBench, 36.3% as of 2026-04, per Kinetic Systems blog. End-to-end task success of computer-use LLM agents on healthcare administration workflows: prior authorizations, denial appeals, and DME ordering; success requires completing every subtask in a task.
Current results
- 1
Claude Opus 4.6 (computer-use agent)36.3%
- 2
GPT-5.4 (computer-use agent)26.7%
- 3
Kimi K2.515.6%
- 4AQwen 3.513.3%
- 5
Gemini 3.1 Pro11.9%
Result detail
| # | model | score | as of | |
|---|---|---|---|---|
| 1 | Claude Opus 4.6 (computer-use agent) Anthropic screenshot-only, detailed prompting; subtask rate ~82% | 36.3% | 2026-04 | |
| 2 | GPT-5.4 (computer-use agent) OpenAI screenshot-only, detailed prompting; subtask rate 82.8% | 26.7% | 2026-04 | |
| 3 | Kimi K2.5 Moonshot AI screenshot-only, detailed prompting | 15.6% | 2026-04 | |
| 4 | A | Qwen 3.5 Alibaba screenshot-only, detailed prompting | 13.3% | 2026-04 |
| 5 | Gemini 3.1 Pro Google screenshot-only, detailed prompting | 11.9% | 2026-04 | |
Scores appear exactly as Kinetic Systems blog publishes them (independently run). Run by Kinetic Systems' research team: independent of the model vendors, though published by a company selling healthcare-admin automation. No GPT-5.6 or Claude 5 rows yet, and no stated refresh cadence.
About the benchmark
| publisher | Kinetic Systems (with Stanford Hospital domain experts) |
|---|---|
| category | agentic and workflow benchmarks |
| released | 2026-04 |
| size | 135 tasks / 1,698 rubric-scored subtasks |
| scale | percentage end-to-end task success 0-100, higher better |
| result basis | independently run |
| source | Kinetic Systems blog |
| last frontier result | 2026-04 |
What is HealthAdminBench?
HealthAdminBench is a agentic and workflow benchmark from Kinetic Systems, released 2026-04: 135 tasks / 1,698 rubric-scored subtasks, scored on a percentage end-to-end task success 0-100 scale. End-to-end task success of computer-use LLM agents on healthcare administration workflows: prior authorizations, denial appeals, and DME ordering; success requires completing every subtask in a task.
Which model leads HealthAdminBench?
Claude Opus 4.6 (computer-use agent) (Anthropic) holds the top current result on HealthAdminBench at 36.3%, per Kinetic Systems blog, as of 2026-04.
Where do the HealthAdminBench numbers come from?
From Kinetic Systems blog (independently run). Run by Kinetic Systems' research team: independent of the model vendors, though published by a company selling healthcare-admin automation. No GPT-5.6 or Claude 5 rows yet, and no stated refresh cadence.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.