GPT-5.5: healthcare benchmark results
OpenAI · 3 boards · updated August 16, 2026
The index currently holds 3 results for GPT-5.5: 56.5 on HealthBench (6 of 8), 70.0% on First, Do NOHARM (v2) (4 of 19), 46.3 ± 1.2 on PhysicianBench (1 of 13). It tops PhysicianBench. Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.
Results by benchmark
| benchmark | score | position | as of |
|---|---|---|---|
| HealthBench via OpenAI Deployment Safety Hub (GPT-5.6 system card + August 2026 updates); benchlm.ai and llm-stats.com mirror | 56.5 | 6 of 8 | 2026-06 |
| First, Do NOHARM (v2) via ARISE MAST technical leaderboard | 70.0% | 4 of 19 | 2026-08 |
| PhysicianBench via PhysicianBench paper (Table 2) | 46.3 ± 1.2 | 1 of 13 | 2026-05 |
Position counts against the source's full board, including rows this index does not mirror. Config caveats, where a source noted any, are on each benchmark's page.
Which healthcare benchmarks is GPT-5.5 scored on?
As of August 16, 2026, GPT-5.5 holds current results on 3 tracked benchmarks: HealthBench, First, Do NOHARM (v2), PhysicianBench.
How does GPT-5.5 rank on them?
GPT-5.5 stands at 56.5 on HealthBench (6 of 8), 70.0% on First, Do NOHARM (v2) (4 of 19), 46.3 ± 1.2 on PhysicianBench (1 of 13). It holds first place on PhysicianBench.
The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.