GPT-5.6 Sol: healthcare benchmark results
OpenAI · 5 boards · updated August 16, 2026
The index currently holds 5 results for GPT-5.6 Sol: 0.605 on HealthBench Professional (2 of 9), 0.331 on HealthBench Hard (2 of 9), 57.0 on HealthBench (4 of 8), 60.2% on MAST (Medical AI Superintelligence Test) (1 of 11), 70.1% on First, Do NOHARM (v2) (3 of 19). It tops MAST (Medical AI Superintelligence Test). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.
Results by benchmark
| benchmark | score | position | as of |
|---|---|---|---|
| HealthBench Professional via healthbenchprofessional.com | 0.605 | 2 of 9 | 2026-08 |
| HealthBench Hard via healthbenchhard.ai | 0.331 | 2 of 9 | 2026-08 |
| HealthBench via OpenAI Deployment Safety Hub (GPT-5.6 system card + August 2026 updates); benchlm.ai and llm-stats.com mirror | 57.0 | 4 of 8 | 2026-06 |
| MAST (Medical AI Superintelligence Test) via ARISE MAST leaderboard | 60.2% | 1 of 11 | 2026-08 |
| First, Do NOHARM (v2) via ARISE MAST technical leaderboard | 70.1% | 3 of 19 | 2026-08 |
Position counts against the source's full board, including rows this index does not mirror. Config caveats, where a source noted any, are on each benchmark's page.
Which healthcare benchmarks is GPT-5.6 Sol scored on?
As of August 16, 2026, GPT-5.6 Sol holds current results on 5 tracked benchmarks: HealthBench Professional, HealthBench Hard, HealthBench, MAST (Medical AI Superintelligence Test), First, Do NOHARM (v2).
How does GPT-5.6 Sol rank on them?
GPT-5.6 Sol stands at 0.605 on HealthBench Professional (2 of 9), 0.331 on HealthBench Hard (2 of 9), 57.0 on HealthBench (4 of 8), 60.2% on MAST (Medical AI Superintelligence Test) (1 of 11), 70.1% on First, Do NOHARM (v2) (3 of 19). It holds first place on MAST (Medical AI Superintelligence Test).
The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.