Clinical Benchmarks

Documentation and coding benchmarks

2 tracked · updated August 16, 2026

The paperwork half of medicine: assigning ICD-10 codes to whole hospital stays and writing SOAP notes that survive a compliance read. The pair of boards here tells one story between them: note generation sits near its ceiling while coding accuracy still lags far behind it.

MedCode (Vals AI)

Vals AI · 2,755 patient records
  1. 1Anthropic logoClaude Opus 563.57%
  2. 2Google logoGemini 3.1 Pro Preview (02/26)59.06%
  3. 3Anthropic logoClaude Fable 556.07%
  4. 4Google logoGemini 3 Flash (12/25)55.92%
  5. 5Google logoGemini 3.5 Flash55.83%
via Vals AI MedCode leaderboard · updated 2026-08full detail

MedScribe (Vals AI)

Vals AI · 100 SOAP-note cases
  1. 1Anthropic logoClaude Opus 590.99%
  2. 2Meta logoMuse Spark 1.2~90
  3. 3Meta logoMuse Spark 1.1~90
  4. 4Anthropic logoClaude Fable 5~90
  5. 5OpenAI logoGPT 5.188.09%
via Vals AI MedScribe leaderboard · updated 2026-08full detail

Which documentation and coding benchmarks have current frontier-model results?

2 as of August 16, 2026: MedCode (Vals AI) (Claude Opus 5 leads at 63.57%); MedScribe (Vals AI) (Claude Opus 5 leads at 90.99%).

The other categories sit on the index: rubric-graded benchmarks, agentic and workflow benchmarks, safety benchmarks, knowledge and exam benchmarks, composite indices.