Composite indices
3 tracked · updated August 16, 2026
Weighted rollups of several underlying evaluations into one number. They are convenient for a first look and dangerous for a last one: each carries its authors' weighting choices, and two composites disagreeing usually means their components differ, not that one is wrong. Read them alongside the single-purpose boards they draw from.
MAST (Medical AI Superintelligence Test)
ARISE AI Research Network · 6-benchmark composite- 1
GPT-5.6 Sol60.2%
- 2
Kimi K360.1%
- 3
Gemini 3.6 Flash59.3%
- 4
Gemini 3.1 Pro58.9%
- 5AQwen3.5 397B A17B57.9%
MedHELM
Stanford CRFM · 121 tasks- 1
Gemini 3.1 Pro (Preview)0.652
- 2
Gemini 3.5 Flash0.642
- 3
Muse Spark (2026-04-08)0.621
- 4
GPT-5.4 mini0.552
- 5
GPT-5.4 (2026-03-05)0.538
Artificial Analysis Healthcare & Medical Index
Artificial Analysis · 4-benchmark compositeWhich composite indices have current frontier-model results?
3 as of August 16, 2026: MAST (Medical AI Superintelligence Test) (GPT-5.6 Sol leads at 60.2%); MedHELM (Gemini 3.1 Pro (Preview) leads at 0.652); Artificial Analysis Healthcare & Medical Index (Claude Opus 5 (Adaptive Reasoning, Max Effort) leads at 51).
The other categories sit on the index: rubric-graded benchmarks, agentic and workflow benchmarks, documentation and coding benchmarks, safety benchmarks, knowledge and exam benchmarks.