benchgap
Vision & documents

MedXpertQA (MM) leaderboard

As of 2026-10-07, the highest measured score on MedXpertQA (MM) is 81.3% by Gemini 3.1 Pro. 114 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Gemini 3.5 Flash81.7%estimated ± 2.4 pp, high confidence
2Gemini 3.1 Pro81.3%measured
3GPT-5.6 Sol80.6%estimated ± 2.4 pp, high confidence
4Claude Mythos 580.4%estimated ± 6.4 pp, low confidence
5Qwen3.8 Max80.4%measured
6Claude Opus 4.7 (Adaptive)79.4%estimated ± 6.4 pp, low confidence
7Claude Opus 4.879.0%estimated ± 6.4 pp, low confidence
8GLM-5.3-Flash78.8%estimated ± 6.4 pp, low confidence
9Qwen3.8-Flash-Next78.5%estimated ± 4.2 pp, medium confidence
10Gemini 3.7 Flash78.5%estimated ± 6.4 pp, low confidence
11Muse Spark78.4%measured
12Muse Spark 1.178.4%estimated ± 6.4 pp, low confidence
13Claude Sonnet 578.4%estimated ± 6.4 pp, low confidence
14Qwen3.8-Omni-Flash78.2%estimated ± 4.2 pp, medium confidence
15Kimi K377.9%estimated ± 2.4 pp, high confidence
16Seed 2.1 Pro77.9%estimated ± 2.4 pp, high confidence
17Claude Opus 5.577.9%estimated ± 6.5 pp, low confidence
18GPT-6 Astra77.8%estimated ± 6.5 pp, low confidence
19GPT-6.1 Sol77.8%estimated ± 6.5 pp, low confidence
20Gemini 3.8 Flash77.8%estimated ± 6.5 pp, low confidence
21Claude Opus 577.8%estimated ± 6.5 pp, low confidence
22Sakana Fugu-Ultra77.7%estimated ± 6.4 pp, low confidence
23Gemini 3.6 Flash77.6%estimated ± 6.5 pp, low confidence
24GPT-6 Sol77.6%estimated ± 6.5 pp, low confidence
25Qwen3.8 Max Preview77.6%estimated ± 6.5 pp, low confidence
26GPT-5.477.1%measured
27Grok 4.577.1%estimated ± 6.5 pp, low confidence
28Sakana Fugu77.1%estimated ± 6.4 pp, low confidence
29GPT-5.577.1%estimated ± 2.4 pp, high confidence
30GPT-6 Luna76.9%estimated ± 6.5 pp, low confidence
31Gemini 3 Pro76.6%estimated ± 2.4 pp, high confidence
32Apodex 1.176.6%estimated ± 6.5 pp, low confidence
33Apodex 1.1 Mini76.6%estimated ± 6.5 pp, low confidence
34Gemini 3.5 Flash-Lite76.5%estimated ± 6.5 pp, low confidence
35Ling 3.0 Flash VL76.5%estimated ± 6.5 pp, low confidence
36Gemini 3 Flash76.3%estimated ± 6.5 pp, low confidence
37Qwen3.8-27B76.3%estimated ± 4.2 pp, medium confidence
38GPT-5.3 Codex76.3%estimated ± 6.5 pp, low confidence
39GPT-5.6 Terra76.0%estimated ± 2.4 pp, high confidence
40DeepSeek V4.1 Flash75.0%estimated ± 6.5 pp, low confidence
41Seed 2.1 Turbo74.8%estimated ± 2.4 pp, high confidence
42Claude Opus 4.774.3%estimated ± 6.5 pp, low confidence
43Mistral Large 474.3%estimated ± 6.5 pp, low confidence
44GPT-5.2-Codex74.2%estimated ± 6.5 pp, low confidence
45Claude Sonnet 4.674.0%estimated ± 6.4 pp, low confidence
46Qwen3.5-122B-A10B73.9%estimated ± 6.4 pp, low confidence
47Nemotron 3 Nano Omni 30B A3B73.5%estimated ± 6.4 pp, low confidence
48GPT-5.273.5%estimated ± 2.4 pp, high confidence
49Kimi K2.673.3%estimated ± 2.4 pp, high confidence
50GPT-5.173.0%estimated ± 6.5 pp, low confidence
51Claude Opus 4.6 (Adaptive)72.8%estimated ± 6.5 pp, low confidence
52Step 3.7 Flash72.6%estimated ± 6.5 pp, low confidence
53dots3-note Preview72.6%estimated ± 2.4 pp, high confidence
54Qwen3.5 397B72.4%estimated ± 2.4 pp, high confidence
55Gemini 3.1 Flash-Lite72.3%estimated ± 6.4 pp, low confidence
56Qwen3.5-27B72.1%estimated ± 6.5 pp, low confidence
57Qwen3.6 Plus72.0%estimated ± 2.4 pp, high confidence
58Gemini 2.5 Pro71.9%estimated ± 6.5 pp, low confidence
59Kimi K2.571.3%estimated ± 2.4 pp, high confidence
60Kimi K2.5 (Reasoning)71.3%estimated ± 2.4 pp, high confidence
61GPT-5.6 Luna71.1%estimated ± 2.4 pp, high confidence
62Qwen3.7 Plus71.0%measured
63GPT-5 (medium)70.5%estimated ± 6.5 pp, low confidence
64Grok 4.370.4%estimated ± 2.4 pp, high confidence
65MiniMax M370.4%estimated ± 2.4 pp, high confidence
66GPT-5 (high)70.2%estimated ± 6.5 pp, low confidence
67Pareto 26.970.2%estimated ± 2.4 pp, high confidence
68MiMo-V2.570.0%estimated ± 2.4 pp, high confidence
69Claude Opus 4.5 Thinking69.7%estimated ± 6.5 pp, low confidence
70Gemma 4 31B67.7%estimated ± 2.4 pp, high confidence
71GPT-5.4 mini67.0%estimated ± 2.4 pp, high confidence
72MiMo-V2.6-Flash66.8%estimated ± 6.5 pp, low confidence
73Grok 4.2065.8%measured
74GLM-5V-Turbo65.7%estimated ± 6.5 pp, low confidence
75Step 5 Preview65.6%estimated ± 2.4 pp, high confidence
76Qwen3.5-35B-A3B65.3%estimated ± 6.5 pp, low confidence
77Qwen3.6-27B65.1%estimated ± 2.4 pp, high confidence
78Claude Opus 4.664.8%measured
79GPT-5.1-Codex64.5%estimated ± 6.5 pp, low confidence
80GPT-5.1-Codex-Max64.5%estimated ± 6.5 pp, low confidence
81Qwen3.6-35B-A3B64.0%estimated ± 2.4 pp, high confidence
82Inkling-Small60.8%estimated ± 2.4 pp, high confidence
83Muse Glimmer 30B60.8%estimated ± 2.4 pp, high confidence
84Gemma 4 26B A4B60.3%estimated ± 2.4 pp, high confidence
85Inkling59.6%estimated ± 2.4 pp, high confidence
86Interfaze Beta53.7%estimated ± 2.4 pp, high confidence
87Claude Opus 4.552.4%estimated ± 2.4 pp, high confidence
88o351.5%estimated ± 6.5 pp, low confidence
89MiMo-V2-Omni50.2%estimated ± 6.5 pp, low confidence
90Gemma 4 12B48.7%measured
91Grok 442.5%estimated ± 6.5 pp, low confidence
92GPT-5.4 nano41.2%estimated ± 2.4 pp, medium confidence
93Claude 4.1 Opus Thinking35.9%estimated ± 6.5 pp, low confidence
94Command A+33.8%estimated ± 2.4 pp, medium confidence
95Gemini 2.5 Flash20.1%estimated ± 6.5 pp, low confidence
96Mistral Medium 3.5 128B16.9%estimated ± 6.5 pp, low confidence
97Grok 4.1 Fast (Reasoning)10.3%estimated ± 6.5 pp, low confidence
98Claude 4 Sonnet7.6%estimated ± 6.5 pp, low confidence
99Llama 4 Maverick6.9%estimated ± 6.5 pp, low confidence
100Grok 4 Fast (Reasoning)6.2%estimated ± 6.5 pp, low confidence
101GPT-4.15.0%estimated ± 6.5 pp, low confidence
102Qwen3-Omni-30B-A3B-Thinking3.5%estimated ± 6.5 pp, low confidence
103GPT-4.1 mini2.0%estimated ± 6.5 pp, low confidence
104Mistral Small 41.0%estimated ± 6.5 pp, low confidence
105Mistral Small 4 (Reasoning)1.0%estimated ± 6.5 pp, low confidence
106LFM2.5-VL-3B0.7%estimated ± 2.4 pp, medium confidence
107Mistral Large 30.7%estimated ± 6.5 pp, low confidence
108Qwen3-Omni-30B-A3B-Instruct0.6%estimated ± 6.5 pp, low confidence
109Gemini 1.5 Pro0.5%estimated ± 6.5 pp, low confidence
110Mistral Medium 30.2%estimated ± 6.5 pp, low confidence
111Llama 4 Scout0.2%estimated ± 6.5 pp, low confidence
112Qwen3.5 397B (Reasoning)0.2%estimated ± 6.5 pp, low confidence
113Gemma 4 E4B0.1%estimated ± 6.5 pp, low confidence
114Grok 4.1 Fast0.0%estimated ± 6.5 pp, low confidence
115Gemma 3 27B0.0%estimated ± 6.5 pp, low confidence
116Gemma 4 E2B0.0%estimated ± 6.5 pp, low confidence
117Nova Pro0.0%estimated ± 6.5 pp, low confidence
118GPT-4o mini0.0%estimated ± 6.5 pp, low confidence
119GPT-4.1 nano0.0%estimated ± 6.5 pp, low confidence
120Claude 3 Haiku0.0%estimated ± 6.5 pp, low confidence
121LFM2.5-VL-1.6B-Extract0.0%estimated ± 6.5 pp, low confidence
122Phi-4 Multimodal Instruct0.0%estimated ± 6.5 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General