First, Do NOHARM (v2): current results
Stanford/Harvard-led consortium (50+ researchers incl. 29 board-certified physicians); hosted by ARISE · 1,100 consultation cases, 10 specialties, 12,747 expert annotations on 4,249 management options · index updated August 16, 2026
Claude Opus 5 holds the top current result on First, Do NOHARM (v2), 74.6% as of 2026-08, per ARISE MAST technical leaderboard. Frequency and severity of potentially harmful errors in LLM-generated medical consultation recommendations (Numerous Options Harm Assessment for Risk in Medicine); primary-care-to-specialist consults.
Current results
- 1
Claude Opus 574.6%
- 2
Kimi K374.0%
- 3
GPT-5.6 Sol70.1%
- 4
GPT-5.570.0%
- 5
Gemini 3.1 Pro62.6%
Result detail
| # | model | score | as of | |
|---|---|---|---|---|
| 1 | Claude Opus 5 Anthropic v2 run on ARISE; 19 models on the board | 74.6% | 2026-08 | |
| 2 | Kimi K3 Moonshot AI | 74.0% | 2026-08 | |
| 3 | GPT-5.6 Sol OpenAI | 70.1% | 2026-08 | |
| 4 | GPT-5.5 OpenAI | 70.0% | 2026-08 | |
| 5 | Gemini 3.1 Pro Google | 62.6% | 2026-08 | |
Scores appear exactly as ARISE MAST technical leaderboard publishes them (official leaderboard). Paper: arXiv 2512.01241. The v1 study found potential for severe harm in up to 24.6 percent of directly applied recommendations, with errors of omission behind more than 80 percent of the severe cases.
About the benchmark
| publisher | Stanford/Harvard-led consortium (50+ researchers incl. 29 board-certified physicians); hosted by ARISE |
|---|---|
| category | safety benchmarks |
| released | 2025-12 |
| size | 1,100 consultation cases, 10 specialties, 12,747 expert annotations on 4,249 management options |
| scale | percentage safety score, higher better |
| result basis | official leaderboard |
| source | ARISE MAST technical leaderboard |
| last frontier result | 2026-08 |
What is First, Do NOHARM (v2)?
First, Do NOHARM (v2) is a safety benchmark from Stanford/Harvard consortium, released 2025-12: 1,100 consultation cases, 10 specialties, 12,747 expert annotations on 4,249 management options, scored on a percentage safety score scale. Frequency and severity of potentially harmful errors in LLM-generated medical consultation recommendations (Numerous Options Harm Assessment for Risk in Medicine); primary-care-to-specialist consults.
Which model leads First, Do NOHARM (v2)?
Claude Opus 5 (Anthropic) holds the top current result on First, Do NOHARM (v2) at 74.6%, per ARISE MAST technical leaderboard, as of 2026-08.
Where do the First, Do NOHARM (v2) numbers come from?
From ARISE MAST technical leaderboard (official leaderboard). Paper: arXiv 2512.01241. The v1 study found potential for severe harm in up to 24.6 percent of directly applied recommendations, with errors of omission behind more than 80 percent of the severe cases.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.