MedXpertQA (MM): current results
TsinghuaC3I (Tsinghua University) · 2,000 multimodal questions (MM subset) · index updated August 16, 2026
Gemini 3.1 Pro holds the top current result on MedXpertQA (MM), 81.3% as of 2026-08, per benchlm.ai mirror of Meta's Muse Spark evaluation. Expert-level multimodal medical multiple-choice QA covering clinical images (X-ray, histology, dermatology, charts) across 17 specialties; MM subset of the 4,460-question MedXpertQA benchmark.
Current results
- 1
Gemini 3.1 Pro81.3%
- 2AQwen3.8 Max80.4%
- 3
Muse Spark78.4%
- 4
GPT-5.477.1%
- 5AQwen3.7 Plus71.0%
- 6
Grok 4.2065.8%
- 7
Claude Opus 4.664.8%
- 8
Gemma 4 12B48.7%
Result detail
| # | model | score | as of | |
|---|---|---|---|---|
| 1 | Gemini 3.1 Pro Google as reported in Meta's Muse Spark eval, mirrored by benchlm | 81.3% | 2026-08 | |
| 2 | A | Qwen3.8 Max Alibaba same source | 80.4% | 2026-08 |
| 3 | Muse Spark Meta same source | 78.4% | 2026-08 | |
| 4 | GPT-5.4 OpenAI same source | 77.1% | 2026-08 | |
| 5 | A | Qwen3.7 Plus Alibaba same source | 71.0% | 2026-08 |
| 6 | Grok 4.20 xAI same source | 65.8% | 2026-08 | |
| 7 | Claude Opus 4.6 Anthropic same source | 64.8% | 2026-08 | |
| 8 | Gemma 4 12B Google same source | 48.7% | 2026-08 | |
Scores appear exactly as benchlm.ai mirror of Meta's Muse Spark evaluation publishes them (vendor-reported scores). The current frontier table is vendor-reported, from Meta's Muse Spark launch evaluation, mirrored display-only on benchlm.ai. The Text subset has no comparable current table.
About the benchmark
| publisher | TsinghuaC3I (Tsinghua University) |
|---|---|
| category | knowledge and exam benchmarks |
| released | 2025-01 |
| size | 2,000 multimodal questions (MM subset) |
| scale | percentage accuracy 0-100, higher better |
| result basis | vendor-reported scores |
| source | benchlm.ai mirror of Meta's Muse Spark evaluation |
| last frontier result | 2026-08 |
What is MedXpertQA (MM)?
MedXpertQA (MM) is a knowledge and exam benchmark from Tsinghua University, released 2025-01: 2,000 multimodal questions (MM subset), scored on a percentage accuracy 0-100 scale. Expert-level multimodal medical multiple-choice QA covering clinical images (X-ray, histology, dermatology, charts) across 17 specialties; MM subset of the 4,460-question MedXpertQA benchmark.
Which model leads MedXpertQA (MM)?
Gemini 3.1 Pro (Google) holds the top current result on MedXpertQA (MM) at 81.3%, per benchlm.ai mirror of Meta's Muse Spark evaluation, as of 2026-08.
Where do the MedXpertQA (MM) numbers come from?
From benchlm.ai mirror of Meta's Muse Spark evaluation (vendor-reported scores). The current frontier table is vendor-reported, from Meta's Muse Spark launch evaluation, mirrored display-only on benchlm.ai. The Text subset has no comparable current table.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.