benchgap
Long context

MRCRv2 leaderboard

As of 2026-10-10, the highest measured score on MRCRv2 is 93.6% by Sakana Fugu-Ultra. 177 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Sakana Fugu-Ultra93.6%measured
2Qwen3.8 Max92.9%measured
3Qwen3.7 Plus91.7%measured
4Qwen3.7 Max90.4%measured
5Sakana Fugu86.6%measured
6Kimi K386.0%estimated ± 11.3 pp, low confidence
7Step 5 Preview86.0%estimated ± 11.3 pp, low confidence
8MiMo-V2.6-Pro86.0%estimated ± 11.3 pp, low confidence
9Claude Fable 5.186.0%estimated ± 11.3 pp, low confidence
10Claude Opus 5.586.0%estimated ± 11.3 pp, low confidence
11GPT-5.586.0%estimated ± 11.3 pp, low confidence
12DeepSeek V4.1 Flash86.0%estimated ± 11.3 pp, low confidence
13GPT-5.6 Sol86.0%estimated ± 11.3 pp, low confidence
14GPT-5.6 Luna85.9%estimated ± 11.3 pp, low confidence
15GPT-6 Sol85.9%estimated ± 11.3 pp, low confidence
16GPT-5.3 Codex85.9%estimated ± 11.3 pp, low confidence
17GPT-6 Luna85.9%estimated ± 11.3 pp, low confidence
18Muse Glimmer 30B85.9%estimated ± 11.3 pp, low confidence
19GPT-5.6 Terra85.9%estimated ± 11.3 pp, low confidence
20GPT-6.1 Sol85.9%estimated ± 11.3 pp, low confidence
21Ling 3.1 Flash85.9%estimated ± 11.3 pp, low confidence
22MiniMax M385.9%estimated ± 11.3 pp, low confidence
23Muse Spark 1.385.9%estimated ± 11.3 pp, low confidence
24Claude Haiku 5.585.9%estimated ± 11.3 pp, low confidence
25Claude Sonnet 5.585.9%estimated ± 11.3 pp, low confidence
26GPT-5.285.9%estimated ± 11.3 pp, low confidence
27Claude Fable 585.9%estimated ± 11.3 pp, low confidence
28GPT-5.2-Codex85.9%estimated ± 11.3 pp, low confidence
29Claude Sonnet 585.9%estimated ± 11.3 pp, low confidence
30Gemini 3.1 Pro85.9%estimated ± 11.3 pp, low confidence
31GPT-5.485.9%estimated ± 11.3 pp, low confidence
32Qwen3.8-27B85.9%estimated ± 11.3 pp, low confidence
33Gemini 3.7 Flash85.9%estimated ± 11.3 pp, low confidence
34Gemini 3.8 Flash85.9%estimated ± 11.3 pp, low confidence
35Mistral Large 485.9%estimated ± 11.3 pp, low confidence
36Kimi K2.685.8%estimated ± 11.3 pp, low confidence
37GPT-6 Astra85.8%estimated ± 11.3 pp, low confidence
38Qwen 3.6 Max (preview)85.8%estimated ± 11.3 pp, low confidence
39DeepSeek V4 Pro 081385.8%estimated ± 11.3 pp, low confidence
40Grok 4.685.8%estimated ± 11.3 pp, low confidence
41Qwen3.8 Max Preview85.8%estimated ± 11.3 pp, low confidence
42Gemini 3.6 Flash85.8%estimated ± 11.3 pp, low confidence
43GLM-5.3-Flash85.8%estimated ± 11.3 pp, low confidence
44GPT-5.185.8%estimated ± 11.3 pp, low confidence
45DeepSeek V4 Flash 073185.7%estimated ± 11.3 pp, low confidence
46Gemini 4 Argon85.7%estimated ± 11.3 pp, low confidence
47GLM-5.385.7%estimated ± 11.3 pp, low confidence
48MiMo-V2.5-Pro85.7%estimated ± 11.3 pp, low confidence
49Qwen3.8-Flash-Next85.7%estimated ± 11.3 pp, low confidence
50Apodex 1.185.7%estimated ± 11.3 pp, low confidence
51Apodex 1.1 Mini85.7%estimated ± 11.3 pp, low confidence
52Claude Opus 585.7%estimated ± 11.3 pp, low confidence
53Grok 4.585.7%estimated ± 11.3 pp, low confidence
54Kimi K2.7 Code85.7%estimated ± 11.3 pp, low confidence
55Hy385.6%estimated ± 11.3 pp, low confidence
56Muse Spark 1.285.6%estimated ± 11.3 pp, low confidence
57Claude Opus 4.7 (Adaptive)85.6%estimated ± 11.3 pp, low confidence
58GLM-5.285.5%estimated ± 11.3 pp, low confidence
59Ling 3.0 Flash VL85.5%estimated ± 11.3 pp, low confidence
60MiniMax M2.785.5%estimated ± 11.3 pp, low confidence
61Qwen3.6 Plus85.5%estimated ± 11.3 pp, low confidence
62GPT-5 (high)85.5%estimated ± 11.3 pp, low confidence
63Claude Opus 4.6 (Adaptive)85.5%estimated ± 11.3 pp, low confidence
64Kimi K2.585.5%estimated ± 11.3 pp, low confidence
65Kimi K2.5 (Reasoning)85.5%estimated ± 11.3 pp, low confidence
66Muse Spark85.5%estimated ± 11.3 pp, low confidence
67Claude Opus 4.885.4%estimated ± 11.3 pp, low confidence
68Muse Spark 1.185.4%estimated ± 11.3 pp, low confidence
69Qwen3.5-27B85.4%estimated ± 11.3 pp, low confidence
70Claude Opus 4.5 Thinking85.3%estimated ± 11.3 pp, low confidence
71Inkling85.3%estimated ± 11.3 pp, low confidence
72Qwen3.6-27B85.3%estimated ± 11.3 pp, low confidence
73GPT-5.4 mini85.2%estimated ± 11.3 pp, low confidence
74Ternary Bonsai 2 27B85.2%estimated ± 11.3 pp, low confidence
75GPT-5.4 nano85.1%estimated ± 11.3 pp, low confidence
76Grok 4.785.1%estimated ± 11.3 pp, low confidence
77Quasar 438B84.9%estimated ± 11.3 pp, low confidence
78Qwen3.5-122B-A10B84.9%estimated ± 11.3 pp, low confidence
79Claude 4.1 Opus Thinking84.8%estimated ± 11.3 pp, low confidence
80Gemini 3 Pro84.8%estimated ± 11.3 pp, low confidence
81GPT-5 (medium)84.8%estimated ± 11.3 pp, low confidence
82Claude Opus 4.784.6%estimated ± 11.3 pp, low confidence
83GLM-584.6%estimated ± 11.3 pp, low confidence
84Inkling-Small84.6%estimated ± 11.3 pp, low confidence
85MiMo-V2-Omni84.2%estimated ± 11.3 pp, low confidence
86o383.9%estimated ± 11.3 pp, low confidence
87MiMo-V2.6-Flash83.6%estimated ± 11.3 pp, low confidence
88Grok 4.1 Fast (Reasoning)83.3%estimated ± 11.3 pp, low confidence
89GLM-5.182.9%estimated ± 11.3 pp, low confidence
90Grok 4 Fast (Reasoning)82.9%estimated ± 11.3 pp, low confidence
91Step 3.7 Flash82.9%estimated ± 11.3 pp, low confidence
92Ling 3.0 Flash82.0%estimated ± 11.3 pp, low confidence
93Ling 3.0 Flash FP882.0%estimated ± 11.3 pp, low confidence
94Qwen3.5-35B-A3B80.1%estimated ± 11.3 pp, low confidence
95GLM-5-Turbo79.5%estimated ± 11.3 pp, low confidence
96Qwen3.6-35B-A3B79.5%estimated ± 11.3 pp, low confidence
97GLM-4.777.6%estimated ± 11.3 pp, low confidence
98Solar Pro 477.6%estimated ± 11.3 pp, low confidence
99Gemini 3.5 Flash77.3%measured
100Claude Opus 4.576.6%estimated ± 11.3 pp, low confidence
101GLM-5V-Turbo75.3%estimated ± 11.3 pp, low confidence
102Gemini 3.5 Flash-Lite72.2%measured
103GPT-5.1-Codex71.1%estimated ± 11.3 pp, low confidence
104GPT-5.1-Codex-Max71.1%estimated ± 11.3 pp, low confidence
105Mistral Medium 3.5 128B71.1%estimated ± 11.3 pp, low confidence
106Gemini 2.5 Pro69.6%estimated ± 11.3 pp, low confidence
107Gemma 4 31B66.4%measured
108Claude Sonnet 4.665.9%estimated ± 11.3 pp, low confidence
109GPT-4.165.9%estimated ± 11.3 pp, low confidence
110MiMo-V2-Pro65.9%estimated ± 11.3 pp, low confidence
111Grok 464.2%estimated ± 11.3 pp, low confidence
112Mercury 2.564.2%estimated ± 11.3 pp, low confidence
113Pokee-Isaac 28B60.7%measured
114Claude Opus 4.658.0%estimated ± 11.3 pp, low confidence
115Nemotron 3 Ultra58.0%estimated ± 11.3 pp, low confidence
116Hy3 Preview56.1%estimated ± 11.3 pp, low confidence
117A.X K251.6%estimated ± 11.3 pp, low confidence
118Nemotron 3 Super 100B49.7%estimated ± 11.3 pp, low confidence
119o145.4%estimated ± 11.3 pp, low confidence
120Gemma 4 26B A4B44.1%measured
121Gemma 4 12B43.4%measured
122Grok 4.341.4%estimated ± 11.3 pp, low confidence
123Qwen3.5 397B41.4%estimated ± 11.3 pp, low confidence
124Qwen3.5 397B (Reasoning)41.4%estimated ± 11.3 pp, low confidence
125Solar Open 232.6%estimated ± 11.3 pp, low confidence
126K-Exaone29.7%estimated ± 11.3 pp, low confidence
127Ling 3.0 Tiny27.5%estimated ± 11.3 pp, low confidence
128MiniCPM5-2B25.7%estimated ± 11.3 pp, low confidence
129Gemma 4 E4B25.4%measured
130DeepSeek V3.1 (Reasoning)24.0%estimated ± 11.3 pp, low confidence
131K-EXAONE 2.023.8%estimated ± 11.3 pp, low confidence
132DeepSeek-R123.6%estimated ± 11.3 pp, low confidence
133Gemini 3 Flash23.5%estimated ± 11.3 pp, low confidence
134Grok Code Fast 123.1%estimated ± 11.3 pp, low confidence
135Kimi K223.1%estimated ± 11.3 pp, low confidence
136Command A+23.1%estimated ± 11.3 pp, low confidence
137GPT-OSS 120B23.0%estimated ± 11.3 pp, low confidence
138Llama 4 Maverick22.9%estimated ± 11.3 pp, low confidence
139Qwen3 Max22.9%estimated ± 11.3 pp, low confidence
140Gemini 2.5 Flash22.9%estimated ± 11.3 pp, low confidence
141Mistral Small 422.9%estimated ± 11.3 pp, low confidence
142Mistral Small 4 (Reasoning)22.9%estimated ± 11.3 pp, low confidence
143GPT-4o22.9%estimated ± 11.3 pp, low confidence
144Nemotron 3.5 Lightning 30B A3B NVFP422.9%estimated ± 11.3 pp, low confidence
145Granite 4.2 30B22.9%estimated ± 11.3 pp, low confidence
146DeepSeek V3.122.8%estimated ± 11.3 pp, low confidence
147GLM-4.5-Air22.8%estimated ± 11.3 pp, low confidence
148DeepSeek V3.222.8%estimated ± 11.3 pp, low confidence
149Granite 4.2 8B22.8%estimated ± 11.3 pp, low confidence
150Claude 4 Sonnet22.8%estimated ± 11.3 pp, low confidence
151GPT-4.1 mini22.8%estimated ± 11.3 pp, low confidence
152DeepSeek V3 032422.8%estimated ± 11.3 pp, low confidence
153Nemotron 3 Nano Omni 30B A3B22.8%estimated ± 11.3 pp, low confidence
154Celeris-122.8%estimated ± 11.3 pp, low confidence
155Nemotron 3 Nano 30B22.8%estimated ± 11.3 pp, low confidence
156Trinity-Large-Preview22.8%estimated ± 11.3 pp, low confidence
157Trinity-Large-Thinking22.8%estimated ± 11.3 pp, low confidence
158North Mini Code22.8%estimated ± 11.3 pp, low confidence
159MiMo-V2-Flash22.8%estimated ± 11.3 pp, low confidence
160Mistral Large 322.8%estimated ± 11.3 pp, low confidence
161GPT-OSS 20B22.8%estimated ± 11.3 pp, low confidence
162Solar Pro 322.8%estimated ± 11.3 pp, low confidence
163Grok 4.1 Fast22.8%estimated ± 11.3 pp, low confidence
164Ling 2.6 Flash22.8%estimated ± 11.3 pp, low confidence
165Mistral Medium 322.8%estimated ± 11.3 pp, low confidence
166DeepSeek V322.8%estimated ± 11.3 pp, low confidence
167Claude 3 Haiku22.8%estimated ± 11.3 pp, low confidence
168Llama 4 Scout22.8%estimated ± 11.3 pp, low confidence
169GLM-4.622.8%estimated ± 11.3 pp, low confidence
170Llama 3.1 405B22.8%estimated ± 11.3 pp, low confidence
171Granite 4.2 3B22.8%estimated ± 11.3 pp, low confidence
172Nova Pro22.8%estimated ± 11.3 pp, low confidence
173GPT-4.1 nano22.8%estimated ± 11.3 pp, low confidence
174Ultravox v0.6 Llama 3.3 70B22.8%estimated ± 11.3 pp, low confidence
175DeepSeek R1 Distill Qwen 32B22.8%estimated ± 11.3 pp, low confidence
176Gemma 3 27B22.8%estimated ± 11.3 pp, low confidence
177Granite-4.0-H-1B22.8%estimated ± 11.3 pp, low confidence
178LFM2.5-2.6B22.8%estimated ± 11.3 pp, low confidence
179Exaone 4.0 1.2B22.8%estimated ± 11.3 pp, low confidence
180Granite-4.0-350M22.8%estimated ± 11.3 pp, low confidence
181Granite-4.0-H-350M22.8%estimated ± 11.3 pp, low confidence
182LFM2.5-8B-A1B22.8%estimated ± 11.3 pp, low confidence
183LFM2.5-VL-1.6B-Extract22.8%estimated ± 11.3 pp, low confidence
184Phi-422.8%estimated ± 11.3 pp, low confidence
185Qwen3-Omni-30B-A3B-Instruct22.8%estimated ± 11.3 pp, low confidence
186Qwen3-Omni-30B-A3B-Thinking22.8%estimated ± 11.3 pp, low confidence
187Sarvam 105B22.8%estimated ± 11.3 pp, low confidence
188Sarvam 30B22.8%estimated ± 11.3 pp, low confidence
189Solar Pro 222.8%estimated ± 11.3 pp, low confidence
190Gemma 4 E2B19.1%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General