benchgap
Knowledge & reasoning

MMMLU leaderboard

As of 2026-10-07, the highest measured score on MMMLU is 90.9% by Interfaze Beta. 62 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1GPT-6 Astra91.6%estimated ± 1.1 pp, low confidence
2Sakana Fugu91.4%estimated ± 1.1 pp, low confidence
3Sakana Fugu-Ultra91.4%estimated ± 1.1 pp, low confidence
4GPT-5.6 Sol91.2%estimated ± 1.1 pp, low confidence
5Gemini 3.1 Pro91.1%estimated ± 1.1 pp, low confidence
6Claude Opus 4.7 (Adaptive)91.1%estimated ± 1.1 pp, low confidence
7Claude Opus 4.890.9%estimated ± 1.1 pp, low confidence
8GPT-5.590.9%estimated ± 1.1 pp, low confidence
9Kimi K390.9%estimated ± 1.1 pp, low confidence
10Step 5 Preview90.9%estimated ± 1.1 pp, low confidence
11Interfaze Beta90.9%measured
12GPT-5.6 Terra90.8%estimated ± 1.1 pp, low confidence
13GPT-5.490.7%estimated ± 1.1 pp, low confidence
14Ornith-1.5-397B90.7%estimated ± 1.1 pp, low confidence
15Gemini 3.5 Flash90.7%estimated ± 1.1 pp, low confidence
16Qwen3.8 Max90.7%estimated ± 1.1 pp, low confidence
17Pareto 26.10 Preview90.6%estimated ± 1.1 pp, medium confidence
18GPT-5.6 Luna90.6%estimated ± 1.1 pp, medium confidence
19Hy4 preview90.6%estimated ± 1.1 pp, medium confidence
20Qwen3.8-Flash-Next90.4%estimated ± 1.1 pp, medium confidence
21Qwen3.7 Max90.3%measured
22GLM-5.290.2%estimated ± 1.1 pp, medium confidence
23Qwen3.8-Omni-Flash90.2%estimated ± 1.1 pp, medium confidence
24DeepSeek V4.1 Flash90.1%estimated ± 1.1 pp, medium confidence
25Kimi K2.690.0%estimated ± 1.1 pp, medium confidence
26Beam90.0%estimated ± 1.1 pp, medium confidence
27DeepSeek V4 Pro 081389.9%estimated ± 1.1 pp, medium confidence
28Inkling-Small89.6%estimated ± 1.1 pp, medium confidence
29Muse Spark89.6%estimated ± 1.1 pp, medium confidence
30Claude Opus 4.689.5%estimated ± 1.1 pp, medium confidence
31Ornith-1.5-35B-A3B89.5%estimated ± 1.1 pp, medium confidence
32Qwen3.8-27B89.5%estimated ± 1.1 pp, medium confidence
33Solar Pro 489.5%estimated ± 1.1 pp, medium confidence
34Grok 4.2089.3%estimated ± 1.1 pp, medium confidence
35DeepSeek V4 Flash 073189.1%estimated ± 1.1 pp, medium confidence
36Inkling89.0%estimated ± 1.1 pp, medium confidence
37Qwen3.7 Plus89.0%measured
38Kimi K2.588.9%estimated ± 1.1 pp, medium confidence
39Hy3 Preview88.7%estimated ± 1.1 pp, medium confidence
40MiniMax M2.788.6%estimated ± 1.1 pp, medium confidence
41Nemotron 3 Ultra88.6%estimated ± 1.1 pp, medium confidence
42Ornith-1.5-9B88.4%estimated ± 1.1 pp, medium confidence
43Solar Open 288.3%estimated ± 1.1 pp, medium confidence
44GLM-5.188.3%estimated ± 1.1 pp, medium confidence
45GLM-588.2%estimated ± 1.1 pp, medium confidence
46Ternary Bonsai 2 27B88.1%estimated ± 1.1 pp, medium confidence
47A.X K288.0%estimated ± 1.1 pp, medium confidence
48Ling 3.0 Flash87.7%estimated ± 1.1 pp, medium confidence
49MAI-Thinking-187.3%estimated ± 1.1 pp, medium confidence
50Ling 3.0 Flash FP887.2%estimated ± 1.1 pp, medium confidence
51K-EXAONE 2.086.6%measured
52Mercury 2.583.9%estimated ± 1.1 pp, medium confidence
53Gemma 4 12B83.4%measured
54Trinity-Large-Thinking81.6%estimated ± 1.1 pp, low confidence
55Nemotron 3.5 Lightning 30B A3B NVFP480.9%estimated ± 1.1 pp, low confidence
56Nemotron 3 Nano Omni 30B A3B77.2%estimated ± 1.1 pp, low confidence
57ZAYA1-8B75.7%estimated ± 1.1 pp, low confidence
58MiniCPM5-2B74.6%estimated ± 1.1 pp, low confidence
59LongCat-Flash-Lite-Sparse73.6%estimated ± 1.1 pp, low confidence
60Trinity-Large-Preview62.8%estimated ± 1.1 pp, low confidence
61Mellum2-12B-A2.5B-Thinking49.8%estimated ± 1.1 pp, low confidence
62ZAYA1-74B-Preview49.1%estimated ± 1.1 pp, low confidence
63LLaDA2.2-mini17.8%estimated ± 1.1 pp, low confidence
64Soofi S 30B-A3B15.9%estimated ± 1.1 pp, low confidence
65Mellum2-12B-A2.5B-Instruct11.7%estimated ± 1.1 pp, low confidence
66MiniCPM5-1B0.9%estimated ± 1.1 pp, low confidence
67LFM2.5-230M0.8%estimated ± 1.1 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General