benchgap
Multilingual

MMLU-ProX leaderboard

As of 2026-10-07, the highest measured score on MMLU-ProX is 87.0% by Qwen3.7 Max. 38 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 5.591.7%estimated ± 1.7 pp, low confidence
2Claude Mythos 591.1%estimated ± 1.7 pp, low confidence
3Claude Sonnet 5.590.4%estimated ± 1.7 pp, low confidence
4Claude Opus 590.1%estimated ± 1.7 pp, low confidence
5Claude Fable 5.190.0%estimated ± 1.7 pp, low confidence
6Claude Opus 4.888.3%estimated ± 1.7 pp, low confidence
7Hy4 preview87.7%estimated ± 1.7 pp, low confidence
8Qwen3.8-Flash-Next87.0%estimated ± 1.7 pp, low confidence
9Qwen3.7 Max87.0%measured
10Qwen3.8-Omni-Flash86.8%estimated ± 1.7 pp, low confidence
11Composer 2.586.6%estimated ± 1.7 pp, low confidence
12Ornith-1.5-397B86.5%estimated ± 1.7 pp, low confidence
13Ornith-1.0-397B86.3%estimated ± 1.7 pp, low confidence
14Laguna S 2.186.1%estimated ± 1.7 pp, low confidence
15Claude Sonnet 586.0%estimated ± 1.7 pp, medium confidence
16Grok 4.585.9%estimated ± 1.7 pp, medium confidence
17Beam85.9%estimated ± 1.7 pp, medium confidence
18SWE-1.785.8%estimated ± 1.7 pp, medium confidence
19Claude Opus 4.585.7%measured
20Kimi K2.685.4%estimated ± 1.7 pp, medium confidence
21Qwen3.7 Plus85.4%measured
22MiniMax M2.785.3%estimated ± 1.7 pp, medium confidence
23DeepSeek V4 Pro 081385.2%estimated ± 1.7 pp, medium confidence
24dots3-note Preview85.0%estimated ± 1.7 pp, medium confidence
25Qwen3.5 397B84.7%measured
26Qwen3.6 Plus84.7%measured
27Composer 284.3%estimated ± 1.7 pp, medium confidence
28DeepSeek V4 Flash 073184.1%estimated ± 1.7 pp, medium confidence
29Ling 3.0 Flash83.8%estimated ± 1.7 pp, medium confidence
30Ornith-1.5-35B-A3B83.4%estimated ± 1.7 pp, medium confidence
31Qwen3.6-27B83.3%estimated ± 1.7 pp, medium confidence
32GLM-583.1%measured
33Nemotron 3 Ultra83.0%measured
34Ornith-1.0-35B82.6%estimated ± 1.7 pp, medium confidence
35Kimi K2.582.3%measured
36Qwen3.5-122B-A10B82.2%measured
37Qwen3.5-27B82.2%measured
38Qwen3.6-35B-A3B81.7%estimated ± 1.7 pp, low confidence
39Qwen3.5-35B-A3B81.0%measured
40Laguna M.180.1%estimated ± 1.7 pp, low confidence
41Laguna XS 2.180.1%estimated ± 1.7 pp, low confidence
42Qwen3 235B 250779.4%measured
43LongCat-Flash-Lite-Sparse78.6%estimated ± 1.7 pp, low confidence
44Laguna XS.277.9%estimated ± 1.7 pp, low confidence
45Ornith-1.5-9B76.5%estimated ± 1.7 pp, low confidence
46Ornith-1.0-9B75.5%estimated ± 1.7 pp, low confidence
47GPT-4.172.7%measured
48Granite 4.2 30B71.1%estimated ± 1.7 pp, low confidence
49DeepSeek V3 032470.5%measured
50Nemotron 3.5 Lightning 30B A3B NVFP468.6%estimated ± 1.7 pp, low confidence
51Granite 4.2 8B65.9%estimated ± 1.7 pp, low confidence
52LLaDA2.2-flash63.2%estimated ± 1.7 pp, low confidence
53GPT-4o61.1%measured
54Phi-449.9%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General