benchgap
Knowledge & reasoning

MedXpertQA (Text) leaderboard

As of 2026-10-07, the highest measured score on MedXpertQA (Text) is 71.5% by Gemini 3.1 Pro. 78 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 5.578.7%estimated ± 8.9 pp, low confidence
2GPT-6 Astra77.7%estimated ± 4.5 pp, low confidence
3Claude Fable 5.176.1%estimated ± 8.9 pp, low confidence
4Sakana Fugu75.1%estimated ± 4.5 pp, low confidence
5Sakana Fugu-Ultra75.1%estimated ± 4.5 pp, low confidence
6Claude Mythos 574.6%estimated ± 8.9 pp, low confidence
7Claude Sonnet 5.572.9%estimated ± 8.9 pp, low confidence
8Claude Opus 572.4%estimated ± 8.9 pp, low confidence
9Gemini 3.1 Pro71.5%measured
10GPT-5.6 Sol70.7%estimated ± 4.5 pp, low confidence
11Muse Spark 1.168.9%estimated ± 8.9 pp, low confidence
12Claude Opus 4.7 (Adaptive)68.9%estimated ± 4.5 pp, medium confidence
13Claude Opus 4.866.3%estimated ± 4.5 pp, medium confidence
14GPT-5.566.3%estimated ± 4.5 pp, medium confidence
15Pareto 26.966.1%estimated ± 8.9 pp, low confidence
16Kimi K365.9%estimated ± 4.5 pp, medium confidence
17Step 5 Preview65.9%estimated ± 4.5 pp, medium confidence
18GPT-5.6 Terra63.5%estimated ± 4.5 pp, medium confidence
19Claude Haiku 5.563.2%estimated ± 8.9 pp, low confidence
20Ornith-1.5-397B63.2%estimated ± 4.5 pp, medium confidence
21Gemini 3.5 Flash62.7%estimated ± 4.5 pp, medium confidence
22Qwen3.8 Max62.4%estimated ± 4.5 pp, medium confidence
23Pareto 26.10 Preview61.7%estimated ± 4.5 pp, medium confidence
24Qwen3.7 Max61.7%estimated ± 4.5 pp, medium confidence
25GPT-5.6 Luna61.3%estimated ± 4.5 pp, medium confidence
26Hy4 preview61.3%estimated ± 4.5 pp, medium confidence
27Claude Sonnet 560.7%estimated ± 8.9 pp, low confidence
28GPT-5.5 Pro60.6%estimated ± 8.9 pp, low confidence
29GPT-5.4 Pro60.2%estimated ± 8.9 pp, low confidence
30GPT-5.459.6%measured
31Qwen3.8-Flash-Next59.2%estimated ± 4.5 pp, medium confidence
32GLM-5.257.5%estimated ± 4.5 pp, medium confidence
33Qwen3.8-Omni-Flash56.8%estimated ± 4.5 pp, medium confidence
34DeepSeek V4.1 Flash56.5%estimated ± 4.5 pp, medium confidence
35Kimi K2.655.3%estimated ± 4.5 pp, medium confidence
36Beam55.3%estimated ± 4.5 pp, medium confidence
37Qwen3.7 Plus54.7%estimated ± 4.5 pp, medium confidence
38DeepSeek V4 Pro 081354.1%estimated ± 4.5 pp, medium confidence
39Interfaze Beta53.5%estimated ± 4.5 pp, medium confidence
40Muse Spark52.6%measured
41Inkling-Small52.3%estimated ± 4.5 pp, medium confidence
42Claude Opus 4.652.1%measured
43Ornith-1.5-35B-A3B51.5%estimated ± 4.5 pp, medium confidence
44Qwen3.8-27B51.5%estimated ± 4.5 pp, medium confidence
45MiMo-V2.5-Pro51.0%estimated ± 8.9 pp, low confidence
46Solar Pro 450.9%estimated ± 4.5 pp, medium confidence
47Grok 4.2050.2%measured
48DeepSeek V4 Flash 073148.6%estimated ± 4.5 pp, low confidence
49Inkling48.1%estimated ± 4.5 pp, low confidence
50Kimi K2.547.4%estimated ± 4.5 pp, low confidence
51Hy3 Preview46.4%estimated ± 4.5 pp, low confidence
52MiniMax M2.745.9%estimated ± 4.5 pp, low confidence
53Nemotron 3 Ultra45.9%estimated ± 4.5 pp, low confidence
54Ornith-1.5-9B44.6%estimated ± 4.5 pp, low confidence
55Solar Open 244.4%estimated ± 4.5 pp, low confidence
56GPT-5.4 mini44.3%estimated ± 8.9 pp, low confidence
57GLM-5.144.2%estimated ± 4.5 pp, low confidence
58GLM-543.7%estimated ± 4.5 pp, low confidence
59Ternary Bonsai 2 27B43.2%estimated ± 4.5 pp, low confidence
60A.X K242.9%estimated ± 4.5 pp, low confidence
61Ling 3.0 Flash41.6%estimated ± 4.5 pp, low confidence
62MAI-Thinking-140.1%estimated ± 4.5 pp, low confidence
63Ling 3.0 Flash FP839.8%estimated ± 4.5 pp, low confidence
64GPT-5.4 nano39.3%estimated ± 8.9 pp, low confidence
65K-EXAONE 2.036.6%estimated ± 4.5 pp, low confidence
66Gemma 4 31B32.8%estimated ± 8.9 pp, low confidence
67Mercury 2.531.9%estimated ± 4.5 pp, low confidence
68Gemma 4 12B31.6%estimated ± 4.5 pp, low confidence
69Trinity-Large-Thinking28.5%estimated ± 4.5 pp, low confidence
70Nemotron 3.5 Lightning 30B A3B NVFP427.7%estimated ± 4.5 pp, low confidence
71Nemotron 3 Nano Omni 30B A3B24.3%estimated ± 4.5 pp, low confidence
72ZAYA1-8B23.2%estimated ± 4.5 pp, low confidence
73MiniCPM5-2B22.5%estimated ± 4.5 pp, low confidence
74LongCat-Flash-Lite-Sparse21.9%estimated ± 4.5 pp, low confidence
75Trinity-Large-Preview17.5%estimated ± 4.5 pp, low confidence
76Gemma 4 26B A4B16.1%estimated ± 8.9 pp, low confidence
77Mellum2-12B-A2.5B-Thinking14.2%estimated ± 4.5 pp, low confidence
78ZAYA1-74B-Preview14.1%estimated ± 4.5 pp, low confidence
79LLaDA2.2-mini8.9%estimated ± 4.5 pp, low confidence
80Soofi S 30B-A3B8.5%estimated ± 4.5 pp, low confidence
81Mellum2-12B-A2.5B-Instruct7.8%estimated ± 4.5 pp, low confidence
82MiniCPM5-1B4.1%estimated ± 4.5 pp, low confidence
83LFM2.5-230M4.0%estimated ± 4.5 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General