benchgap
Vision & documents

ERQA leaderboard

As of 2026-10-07, the highest measured score on ERQA is 77.8% by Qwen3.8 Max. 113 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Qwen3.8 Max77.8%measured
2Kimi K375.8%estimated ± 2.4 pp, medium confidence
3Claude Mythos 573.0%estimated ± 3.7 pp, high confidence
4Qwen3.8-Flash-Next72.3%measured
5Seed 2.1 Pro72.0%measured
6Claude Opus 4.7 (Adaptive)71.8%estimated ± 3.7 pp, high confidence
7Seed 2.1 Turbo71.3%measured
8Claude Opus 4.871.2%estimated ± 3.7 pp, high confidence
9Qwen3.8-Omni-Flash71.0%measured
10GLM-5.3-Flash71.0%estimated ± 3.7 pp, high confidence
11Gemini 3.7 Flash70.6%estimated ± 3.7 pp, high confidence
12Muse Spark 1.170.4%estimated ± 3.7 pp, high confidence
13Claude Sonnet 570.4%estimated ± 3.7 pp, high confidence
14Qwen3.7 Plus69.8%measured
15Sakana Fugu-Ultra69.5%estimated ± 3.7 pp, high confidence
16Gemini 3.1 Pro69.4%measured
17Claude Opus 5.568.9%estimated ± 3.9 pp, medium confidence
18GPT-6 Astra68.9%estimated ± 3.9 pp, medium confidence
19GPT-6.1 Sol68.9%estimated ± 3.9 pp, medium confidence
20Gemini 3.8 Flash68.9%estimated ± 3.9 pp, medium confidence
21Claude Opus 568.9%estimated ± 3.9 pp, medium confidence
22GPT-5.6 Sol68.8%estimated ± 3.9 pp, medium confidence
23Gemini 3.6 Flash68.8%estimated ± 3.9 pp, medium confidence
24GPT-6 Sol68.8%estimated ± 3.9 pp, medium confidence
25Qwen3.8 Max Preview68.8%estimated ± 3.9 pp, medium confidence
26Sakana Fugu68.7%estimated ± 3.7 pp, high confidence
27GPT-5.6 Terra68.5%estimated ± 3.9 pp, high confidence
28Grok 4.568.5%estimated ± 3.9 pp, high confidence
29GPT-5.568.4%estimated ± 3.9 pp, high confidence
30GPT-6 Luna68.3%estimated ± 3.9 pp, high confidence
31Gemini 3.5 Flash68.2%estimated ± 3.7 pp, high confidence
32Apodex 1.168.1%estimated ± 3.9 pp, high confidence
33Apodex 1.1 Mini68.1%estimated ± 3.9 pp, high confidence
34Gemini 3.5 Flash-Lite68.1%estimated ± 3.9 pp, high confidence
35Ling 3.0 Flash VL68.1%estimated ± 3.9 pp, high confidence
36Gemini 3 Flash67.9%estimated ± 3.9 pp, high confidence
37GPT-5.6 Luna67.9%estimated ± 3.9 pp, high confidence
38MiniMax M367.9%estimated ± 3.9 pp, high confidence
39GPT-5.3 Codex67.8%estimated ± 3.9 pp, high confidence
40Grok 4.367.6%estimated ± 3.9 pp, high confidence
41Inkling67.1%estimated ± 3.7 pp, high confidence
42Qwen3.5 397B66.9%estimated ± 2.4 pp, low confidence
43Inkling-Small66.7%estimated ± 3.7 pp, high confidence
44Qwen3.6-35B-A3B66.7%estimated ± 3.2 pp, medium confidence
45DeepSeek V4.1 Flash66.6%estimated ± 3.9 pp, high confidence
46MiMo-V2.566.5%estimated ± 3.7 pp, high confidence
47Qwen3.6 Plus66.0%estimated ± 2.4 pp, low confidence
48Claude Opus 4.765.9%estimated ± 3.9 pp, high confidence
49Mistral Large 465.9%estimated ± 3.9 pp, high confidence
50Step 5 Preview65.9%estimated ± 3.9 pp, high confidence
51GPT-5.2-Codex65.7%estimated ± 3.9 pp, high confidence
52dots3-note Preview65.6%estimated ± 2.4 pp, low confidence
53Qwen3.8-27B65.5%measured
54GPT-5.465.4%measured
55Muse Glimmer 30B65.3%estimated ± 3.7 pp, high confidence
56Kimi K2.665.2%estimated ± 2.4 pp, low confidence
57Muse Spark64.7%measured
58Claude Sonnet 4.664.5%estimated ± 3.7 pp, high confidence
59GPT-5.164.2%estimated ± 3.9 pp, high confidence
60Claude Opus 4.6 (Adaptive)64.0%estimated ± 3.9 pp, high confidence
61Kimi K2.564.0%estimated ± 3.9 pp, high confidence
62Kimi K2.5 (Reasoning)64.0%estimated ± 3.9 pp, high confidence
63Gemini 3 Pro64.0%estimated ± 2.4 pp, low confidence
64Nemotron 3 Nano Omni 30B A3B63.9%estimated ± 3.7 pp, high confidence
65Step 3.7 Flash63.8%estimated ± 3.9 pp, high confidence
66Qwen3.5-122B-A10B63.5%estimated ± 2.4 pp, low confidence
67Qwen3.5-27B63.2%estimated ± 2.4 pp, low confidence
68Gemini 2.5 Pro62.7%estimated ± 3.9 pp, high confidence
69Qwen3.6-27B62.5%measured
70Pareto 26.962.4%estimated ± 6.1 pp, medium confidence
71Gemini 3.1 Flash-Lite62.1%estimated ± 3.7 pp, high confidence
72GPT-5 (medium)60.8%estimated ± 3.9 pp, high confidence
73GPT-5 (high)60.4%estimated ± 3.9 pp, high confidence
74Qwen3.5-35B-A3B60.3%estimated ± 2.4 pp, low confidence
75Claude Opus 4.5 Thinking59.7%estimated ± 3.9 pp, high confidence
76GPT-5.259.1%estimated ± 2.4 pp, low confidence
77Gemma 4 31B57.0%estimated ± 3.9 pp, high confidence
78GPT-5.4 mini56.4%estimated ± 3.9 pp, high confidence
79Ternary Bonsai 2 27B55.4%estimated ± 3.2 pp, low confidence
80MiMo-V2.6-Flash55.4%estimated ± 3.9 pp, high confidence
81Gemma 4 12B54.9%estimated ± 2.4 pp, low confidence
82Grok 4.2054.1%measured
83GLM-5V-Turbo53.7%estimated ± 3.9 pp, high confidence
84GPT-5.1-Codex51.8%estimated ± 3.9 pp, high confidence
85GPT-5.1-Codex-Max51.8%estimated ± 3.9 pp, high confidence
86Claude Opus 4.651.6%measured
87Command A+49.0%estimated ± 3.7 pp, medium confidence
88Claude Opus 4.548.6%estimated ± 2.4 pp, low confidence
89Interfaze Beta46.6%estimated ± 6.1 pp, low confidence
90LFM2.5-VL-3B39.7%estimated ± 3.2 pp, low confidence
91o332.6%estimated ± 3.9 pp, medium confidence
92MiMo-V2-Omni30.9%estimated ± 3.9 pp, medium confidence
93Gemma 4 26B A4B25.0%estimated ± 3.9 pp, medium confidence
94ZAYA1-VL-8B23.6%estimated ± 3.2 pp, low confidence
95Grok 421.9%estimated ± 3.9 pp, medium confidence
96North Micro Vision Instruct19.0%estimated ± 3.2 pp, low confidence
97Claude 4.1 Opus Thinking15.7%estimated ± 3.9 pp, medium confidence
98LFM2.5-VL-450M13.7%estimated ± 3.2 pp, low confidence
99Gemini 2.5 Flash5.6%estimated ± 3.9 pp, medium confidence
100GPT-5.4 nano5.3%estimated ± 3.9 pp, medium confidence
101Mistral Medium 3.5 128B4.2%estimated ± 3.9 pp, medium confidence
102Grok 4.1 Fast (Reasoning)1.9%estimated ± 3.9 pp, medium confidence
103Claude 4 Sonnet1.2%estimated ± 3.9 pp, medium confidence
104Llama 4 Maverick1.1%estimated ± 3.9 pp, medium confidence
105Grok 4 Fast (Reasoning)0.9%estimated ± 3.9 pp, medium confidence
106GPT-4.10.7%estimated ± 3.9 pp, medium confidence
107Qwen3-Omni-30B-A3B-Thinking0.4%estimated ± 3.9 pp, medium confidence
108GPT-4.1 mini0.2%estimated ± 3.9 pp, medium confidence
109Mistral Small 40.1%estimated ± 3.9 pp, medium confidence
110Mistral Small 4 (Reasoning)0.1%estimated ± 3.9 pp, medium confidence
111Mistral Large 30.0%estimated ± 3.9 pp, medium confidence
112Qwen3-Omni-30B-A3B-Instruct0.0%estimated ± 3.9 pp, medium confidence
113Gemini 1.5 Pro0.0%estimated ± 3.9 pp, medium confidence
114Mistral Medium 30.0%estimated ± 3.9 pp, medium confidence
115Llama 4 Scout0.0%estimated ± 3.9 pp, medium confidence
116Qwen3.5 397B (Reasoning)0.0%estimated ± 3.9 pp, medium confidence
117Gemma 4 E4B0.0%estimated ± 3.9 pp, medium confidence
118Grok 4.1 Fast0.0%estimated ± 3.9 pp, medium confidence
119Gemma 3 27B0.0%estimated ± 3.9 pp, medium confidence
120Gemma 4 E2B0.0%estimated ± 3.9 pp, medium confidence
121Nova Pro0.0%estimated ± 3.9 pp, medium confidence
122Claude 3 Haiku0.0%estimated ± 3.9 pp, medium confidence
123GPT-4.1 nano0.0%estimated ± 3.9 pp, medium confidence
124GPT-4o mini0.0%estimated ± 3.9 pp, medium confidence
125LFM2.5-VL-1.6B-Extract0.0%estimated ± 3.9 pp, medium confidence
126Phi-4 Multimodal Instruct0.0%estimated ± 3.9 pp, medium confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General