MedHELM: current results
Stanford CRFM / HAI and multi-institution collaborators · 121 tasks / 31 datasets · index updated August 16, 2026
Gemini 3.1 Pro (Preview) holds the top current result on MedHELM, 0.652 as of 2026-05, per MedHELM leaderboard (medhelm.org), v5.0.0. Holistic evaluation of LLMs on 121 clinical tasks across 5 categories and 22 subcategories (31 datasets) in a clinician-validated taxonomy; ranked by mean win rate.
Current results
- 1
Gemini 3.1 Pro (Preview)0.652
- 2
Gemini 3.5 Flash0.642
- 3
Muse Spark (2026-04-08)0.621
- 4
GPT-5.4 mini0.552
- 5
GPT-5.4 (2026-03-05)0.538
- 6
Gemini 2.5 Pro0.529
- 7DDeepSeek R10.485
- 8
Claude 4.6 Opus0.456
- 9
Claude 3.7 Sonnet0.45
- 10
Gemini 2.0 Flash0.342
Result detail
| # | model | score | as of | |
|---|---|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) Google | 0.652 | 2026-05 | |
| 2 | Gemini 3.5 Flash Google | 0.642 | 2026-05 | |
| 3 | Muse Spark (2026-04-08) Meta | 0.621 | 2026-05 | |
| 4 | GPT-5.4 mini OpenAI | 0.552 | 2026-05 | |
| 5 | GPT-5.4 (2026-03-05) OpenAI | 0.538 | 2026-05 | |
| 6 | Gemini 2.5 Pro Google | 0.529 | 2026-05 | |
| 7 | D | DeepSeek R1 DeepSeek | 0.485 | 2026-05 |
| 8 | Claude 4.6 Opus Anthropic | 0.456 | 2026-05 | |
| 9 | Claude 3.7 Sonnet Anthropic | 0.45 | 2026-05 | |
| 10 | Gemini 2.0 Flash Google | 0.342 | 2026-05 | |
Scores appear exactly as MedHELM leaderboard (medhelm.org), v5.0.0 publishes them (official leaderboard). Version 5.0.0, last updated May 14, 2026, run by the Stanford-led maintainers on a roughly quarterly cadence. No Claude 5 family or GPT-5.6 rows yet. Mean win rate is relative to the evaluated cohort, so scores shift whenever the model set changes.
About the benchmark
| publisher | Stanford CRFM / HAI and multi-institution collaborators |
|---|---|
| category | composite indices |
| released | 2025-02 |
| size | 121 tasks / 31 datasets |
| scale | mean win rate 0-1, higher better |
| result basis | official leaderboard |
| source | MedHELM leaderboard (medhelm.org), v5.0.0 |
| last frontier result | 2026-05 |
What is MedHELM?
MedHELM is a composite benchmark from Stanford CRFM, released 2025-02: 121 tasks / 31 datasets, scored on a mean win rate 0-1 scale. Holistic evaluation of LLMs on 121 clinical tasks across 5 categories and 22 subcategories (31 datasets) in a clinician-validated taxonomy; ranked by mean win rate.
Which model leads MedHELM?
Gemini 3.1 Pro (Preview) (Google) holds the top current result on MedHELM at 0.652, per MedHELM leaderboard (medhelm.org), v5.0.0, as of 2026-05.
Where do the MedHELM numbers come from?
From MedHELM leaderboard (medhelm.org), v5.0.0 (official leaderboard). Version 5.0.0, last updated May 14, 2026, run by the Stanford-led maintainers on a roughly quarterly cadence. No Claude 5 family or GPT-5.6 rows yet. Mean win rate is relative to the evaluated cohort, so scores shift whenever the model set changes.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.