Clinical Benchmarks

Composite indices

3 tracked · updated August 16, 2026

Weighted rollups of several underlying evaluations into one number. They are convenient for a first look and dangerous for a last one: each carries its authors' weighting choices, and two composites disagreeing usually means their components differ, not that one is wrong. Read them alongside the single-purpose boards they draw from.

MAST (Medical AI Superintelligence Test)

ARISE AI Research Network · 6-benchmark composite
  1. 1OpenAI logoGPT-5.6 Sol60.2%
  2. 2Moonshot AI logoKimi K360.1%
  3. 3Google logoGemini 3.6 Flash59.3%
  4. 4Google logoGemini 3.1 Pro58.9%
  5. 5AQwen3.5 397B A17B57.9%
via ARISE MAST leaderboard · updated 2026-08full detail

MedHELM

Stanford CRFM · 121 tasks
  1. 1Google logoGemini 3.1 Pro (Preview)0.652
  2. 2Google logoGemini 3.5 Flash0.642
  3. 3Meta logoMuse Spark (2026-04-08)0.621
  4. 4OpenAI logoGPT-5.4 mini0.552
  5. 5OpenAI logoGPT-5.4 (2026-03-05)0.538

Artificial Analysis Healthcare & Medical Index

Artificial Analysis · 4-benchmark composite
  1. 1Anthropic logoClaude Opus 5 (Adaptive Reasoning, Max Effort)51
  2. 2Anthropic logoClaude Opus 5 (Adaptive Reasoning, Xhigh Effort)51
  3. 3Anthropic logoClaude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)51

Which composite indices have current frontier-model results?

3 as of August 16, 2026: MAST (Medical AI Superintelligence Test) (GPT-5.6 Sol leads at 60.2%); MedHELM (Gemini 3.1 Pro (Preview) leads at 0.652); Artificial Analysis Healthcare & Medical Index (Claude Opus 5 (Adaptive Reasoning, Max Effort) leads at 51).

The other categories sit on the index: rubric-graded benchmarks, agentic and workflow benchmarks, documentation and coding benchmarks, safety benchmarks, knowledge and exam benchmarks.