Documentation and coding benchmarks
2 tracked · updated August 16, 2026
The paperwork half of medicine: assigning ICD-10 codes to whole hospital stays and writing SOAP notes that survive a compliance read. The pair of boards here tells one story between them: note generation sits near its ceiling while coding accuracy still lags far behind it.
MedCode (Vals AI)
Vals AI · 2,755 patient records- 1
Claude Opus 563.57%
- 2
Gemini 3.1 Pro Preview (02/26)59.06%
- 3
Claude Fable 556.07%
- 4
Gemini 3 Flash (12/25)55.92%
- 5
Gemini 3.5 Flash55.83%
MedScribe (Vals AI)
Vals AI · 100 SOAP-note cases- 1
Claude Opus 590.99%
- 2
Muse Spark 1.2~90
- 3
Muse Spark 1.1~90
- 4
Claude Fable 5~90
- 5
GPT 5.188.09%
Which documentation and coding benchmarks have current frontier-model results?
2 as of August 16, 2026: MedCode (Vals AI) (Claude Opus 5 leads at 63.57%); MedScribe (Vals AI) (Claude Opus 5 leads at 90.99%).
The other categories sit on the index: rubric-graded benchmarks, agentic and workflow benchmarks, safety benchmarks, knowledge and exam benchmarks, composite indices.