Clinical Benchmarks

HealthAdminBench: current results

Kinetic Systems (with Stanford Hospital domain experts) · 135 tasks / 1,698 rubric-scored subtasks · index updated August 16, 2026

Claude Opus 4.6 (computer-use agent) holds the top current result on HealthAdminBench, 36.3% as of 2026-04, per Kinetic Systems blog. End-to-end task success of computer-use LLM agents on healthcare administration workflows: prior authorizations, denial appeals, and DME ordering; success requires completing every subtask in a task.

Current results

Result detail

#modelscoreas of
1Anthropic logoClaude Opus 4.6 (computer-use agent) Anthropic
screenshot-only, detailed prompting; subtask rate ~82%
36.3%2026-04
2OpenAI logoGPT-5.4 (computer-use agent) OpenAI
screenshot-only, detailed prompting; subtask rate 82.8%
26.7%2026-04
3Moonshot AI logoKimi K2.5 Moonshot AI
screenshot-only, detailed prompting
15.6%2026-04
4AQwen 3.5 Alibaba
screenshot-only, detailed prompting
13.3%2026-04
5Google logoGemini 3.1 Pro Google
screenshot-only, detailed prompting
11.9%2026-04

Scores appear exactly as Kinetic Systems blog publishes them (independently run). Run by Kinetic Systems' research team: independent of the model vendors, though published by a company selling healthcare-admin automation. No GPT-5.6 or Claude 5 rows yet, and no stated refresh cadence.

About the benchmark

publisherKinetic Systems (with Stanford Hospital domain experts)
categoryagentic and workflow benchmarks
released2026-04
size135 tasks / 1,698 rubric-scored subtasks
scalepercentage end-to-end task success 0-100, higher better
result basisindependently run
sourceKinetic Systems blog
last frontier result2026-04

What is HealthAdminBench?

HealthAdminBench is a agentic and workflow benchmark from Kinetic Systems, released 2026-04: 135 tasks / 1,698 rubric-scored subtasks, scored on a percentage end-to-end task success 0-100 scale. End-to-end task success of computer-use LLM agents on healthcare administration workflows: prior authorizations, denial appeals, and DME ordering; success requires completing every subtask in a task.

Which model leads HealthAdminBench?

Claude Opus 4.6 (computer-use agent) (Anthropic) holds the top current result on HealthAdminBench at 36.3%, per Kinetic Systems blog, as of 2026-04.

Where do the HealthAdminBench numbers come from?

From Kinetic Systems blog (independently run). Run by Kinetic Systems' research team: independent of the model vendors, though published by a company selling healthcare-admin automation. No GPT-5.6 or Claude 5 rows yet, and no stated refresh cadence.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.