CHI-Bench: current results
actAVA.ai · 75 workflows (25 per domain), 21 healthcare applications, 200+ MCP tools · index updated August 16, 2026
erius + claude-opus-5 holds the top current result on CHI-Bench, 54.7% as of 2026-08, per CHI-Bench leaderboard (actAVA). Long-horizon US healthcare operations workflows for agents (prior authorization, utilization management, care management), 60-80 step tasks across 4-6 stages, judged by deterministic unit tests plus an LLM judge for evidence grounding, consent, and cross-stage consistency.
Current results
- 1
erius + claude-opus-554.7%
- 2
erius + claude-opus-4-837.3%
- 3
claude-code + claude-opus-537.3%
- 4
claude-code + claude-opus-4-833.3%
- 5
claude-code + claude-opus-4-628.0%
- 6
claude-code + claude-sonnet-4-626.2%
- 7
codex + gpt-5.6-sol25.3%
- 8
openai-agents + kimi-k325.3%
- 9
claude-code + claude-opus-4-724.4%
- 10
claude-code + claude-fable-524.0%
Result detail
| # | model | score | as of | |
|---|---|---|---|---|
| 1 | erius + claude-opus-5 Humana (harness) / Anthropic (model) community-submitted harness config validated by automated workspace judge | 54.7% | 2026-08 | |
| 2 | erius + claude-opus-4-8 Humana / Anthropic | 37.3% | 2026-08 | |
| 3 | claude-code + claude-opus-5 Anthropic | 37.3% | 2026-08 | |
| 4 | claude-code + claude-opus-4-8 Anthropic | 33.3% | 2026-08 | |
| 5 | claude-code + claude-opus-4-6 Anthropic | 28.0% | 2026-08 | |
| 6 | claude-code + claude-sonnet-4-6 Anthropic | 26.2% | 2026-08 | |
| 7 | codex + gpt-5.6-sol OpenAI | 25.3% | 2026-08 | |
| 8 | openai-agents + kimi-k3 Moonshot AI | 25.3% | 2026-08 | |
| 9 | claude-code + claude-opus-4-7 Anthropic | 24.4% | 2026-08 | |
| 10 | claude-code + claude-fable-5 Anthropic | 24.0% | 2026-08 | |
Scores appear exactly as CHI-Bench leaderboard (actAVA) publishes them (mixed sources). Released May 20, 2026; updated August 12, 2026 across 45 harness configurations. The launch report led with reliability, not capability: no agent stayed above 20 percent across three identical runs. Harness choice matters as much as model choice, and the board accepts community submissions, so rows mix author-run and submitted results.
About the benchmark
| publisher | actAVA.ai |
|---|---|
| category | agentic and workflow benchmarks |
| released | 2026-05 |
| size | 75 workflows (25 per domain), 21 healthcare applications, 200+ MCP tools |
| scale | pass@1 with binary 0/1 reward, higher better |
| result basis | mixed sources |
| source | CHI-Bench leaderboard (actAVA) |
| last frontier result | 2026-08 |
What is CHI-Bench?
CHI-Bench is a agentic and workflow benchmark from actAVA, released 2026-05: 75 workflows (25 per domain), 21 healthcare applications, 200+ MCP tools, scored on a pass@1 with binary 0/1 reward scale. Long-horizon US healthcare operations workflows for agents (prior authorization, utilization management, care management), 60-80 step tasks across 4-6 stages, judged by deterministic unit tests plus an LLM judge for evidence grounding, consent, and cross-stage consistency.
Which model leads CHI-Bench?
erius + claude-opus-5 (Humana (harness) / Anthropic (model)) holds the top current result on CHI-Bench at 54.7%, per CHI-Bench leaderboard (actAVA), as of 2026-08.
Where do the CHI-Bench numbers come from?
From CHI-Bench leaderboard (actAVA) (mixed sources). Released May 20, 2026; updated August 12, 2026 across 45 harness configurations. The launch report led with reliability, not capability: no agent stayed above 20 percent across three identical runs. Harness choice matters as much as model choice, and the board accepts community submissions, so rows mix author-run and submitted results.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.