Codex (GPT-5.6-sol): healthcare benchmark results
OpenAI · 2 boards · updated August 16, 2026
The index currently holds 2 results for Codex (GPT-5.6-sol): 45% on HealthAgentBench (2 of 12), 25.3% on CHI-Bench (7 of 45). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.
Results by benchmark
| benchmark | score | position | as of |
|---|---|---|---|
| HealthAgentBench via HealthAgentBench leaderboard (Microsoft GitHub Pages) | 45% | 2 of 12 | 2026-07 |
| CHI-Bench via CHI-Bench leaderboard (actAVA) | 25.3% | 7 of 45 | 2026-08 |
Position counts against the source's full board, including rows this index does not mirror. Config caveats, where a source noted any, are on each benchmark's page. Sources also list this model as "codex + gpt-5.6-sol".
Which healthcare benchmarks is Codex (GPT-5.6-sol) scored on?
As of August 16, 2026, Codex (GPT-5.6-sol) holds current results on 2 tracked benchmarks: HealthAgentBench, CHI-Bench.
How does Codex (GPT-5.6-sol) rank on them?
Codex (GPT-5.6-sol) stands at 45% on HealthAgentBench (2 of 12), 25.3% on CHI-Bench (7 of 45).
The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.