benchgap
Vision & documents

RealWorldQA leaderboard

As of 2026-10-07, the highest measured score on RealWorldQA is 88.5% by Qwen3.8-Flash-Next. 114 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 5.591.4%estimated ± 1.2 pp, low confidence
2GPT-6 Astra91.0%estimated ± 1.2 pp, low confidence
3GPT-6.1 Sol90.6%estimated ± 1.2 pp, low confidence
4Gemini 3.8 Flash90.4%estimated ± 1.2 pp, low confidence
5Claude Opus 590.0%estimated ± 1.2 pp, low confidence
6GPT-5.6 Sol89.3%estimated ± 1.2 pp, low confidence
7Gemini 3.6 Flash89.2%estimated ± 1.2 pp, low confidence
8GPT-6 Sol89.1%estimated ± 1.2 pp, low confidence
9Qwen3.8 Max Preview89.0%estimated ± 1.2 pp, low confidence
10Qwen3.8-Flash-Next88.5%measured
11Kimi K388.2%estimated ± 0.8 pp, medium confidence
12Seed 2.1 Pro88.2%estimated ± 0.8 pp, medium confidence
13Claude Mythos 588.0%estimated ± 1.1 pp, medium confidence
14Qwen3.8 Max88.0%measured
15GPT-5.6 Terra88.0%estimated ± 1.2 pp, low confidence
16Grok 4.587.8%estimated ± 1.2 pp, medium confidence
17Qwen3.8-Omni-Flash87.7%measured
18GPT-5.587.6%estimated ± 1.2 pp, medium confidence
19Claude Opus 4.7 (Adaptive)87.5%estimated ± 1.1 pp, medium confidence
20GPT-6 Luna87.5%estimated ± 1.2 pp, medium confidence
21Claude Opus 4.887.3%estimated ± 1.1 pp, medium confidence
22Apodex 1.187.2%estimated ± 1.2 pp, medium confidence
23Apodex 1.1 Mini87.2%estimated ± 1.2 pp, medium confidence
24GLM-5.3-Flash87.2%estimated ± 1.1 pp, medium confidence
25Gemini 3.5 Flash-Lite87.1%estimated ± 1.2 pp, medium confidence
26Ling 3.0 Flash VL87.1%estimated ± 1.2 pp, medium confidence
27Gemini 3.1 Pro87.0%estimated ± 1.1 pp, medium confidence
28Gemini 3.7 Flash87.0%estimated ± 1.1 pp, medium confidence
29Muse Spark 1.187.0%estimated ± 1.1 pp, medium confidence
30Claude Sonnet 586.9%estimated ± 1.1 pp, medium confidence
31Qwen3.7 Plus86.9%measured
32Gemini 3 Flash86.9%estimated ± 1.2 pp, medium confidence
33GPT-5.6 Luna86.9%estimated ± 1.2 pp, medium confidence
34MiniMax M386.9%estimated ± 1.2 pp, medium confidence
35GPT-5.3 Codex86.8%estimated ± 1.2 pp, medium confidence
36Grok 4.386.6%estimated ± 1.2 pp, medium confidence
37Sakana Fugu-Ultra86.6%estimated ± 1.1 pp, medium confidence
38Seed 2.1 Turbo86.5%estimated ± 0.8 pp, medium confidence
39Sakana Fugu86.3%estimated ± 1.1 pp, medium confidence
40Pareto 26.986.1%estimated ± 4.8 pp, medium confidence
41Gemini 3.5 Flash86.1%estimated ± 1.1 pp, medium confidence
42DeepSeek V4.1 Flash86.0%estimated ± 1.2 pp, medium confidence
43Qwen3.8-27B85.9%measured
44Claude Opus 4.785.7%estimated ± 1.2 pp, medium confidence
45Mistral Large 485.7%estimated ± 1.2 pp, medium confidence
46Step 5 Preview85.7%estimated ± 1.2 pp, medium confidence
47GPT-5.485.7%estimated ± 1.1 pp, medium confidence
48GPT-5.2-Codex85.7%estimated ± 1.2 pp, medium confidence
49Inkling85.6%estimated ± 1.1 pp, medium confidence
50Inkling-Small85.4%estimated ± 1.1 pp, medium confidence
51Muse Spark85.4%estimated ± 1.1 pp, medium confidence
52MiMo-V2.585.4%estimated ± 1.1 pp, medium confidence
53Qwen3.6-35B-A3B85.3%measured
54GPT-5.185.2%estimated ± 1.2 pp, medium confidence
55Claude Opus 4.6 (Adaptive)85.2%estimated ± 1.2 pp, medium confidence
56Kimi K2.585.2%estimated ± 1.2 pp, medium confidence
57Kimi K2.5 (Reasoning)85.2%estimated ± 1.2 pp, medium confidence
58Step 3.7 Flash85.1%estimated ± 1.2 pp, medium confidence
59Muse Glimmer 30B84.9%estimated ± 1.1 pp, medium confidence
60Gemini 2.5 Pro84.9%estimated ± 1.2 pp, medium confidence
61Command A+84.7%estimated ± 0.8 pp, medium confidence
62Nemotron 3 Nano Omni 30B A3B84.7%estimated ± 0.8 pp, medium confidence
63Claude Sonnet 4.684.6%estimated ± 1.1 pp, low confidence
64GPT-5 (medium)84.6%estimated ± 1.2 pp, low confidence
65GPT-5 (high)84.5%estimated ± 1.2 pp, low confidence
66Claude Opus 4.5 Thinking84.4%estimated ± 1.2 pp, low confidence
67Interfaze Beta84.2%estimated ± 4.8 pp, medium confidence
68Qwen3.6-27B84.1%measured
69Gemma 4 31B84.1%estimated ± 1.2 pp, low confidence
70GPT-5.4 mini84.0%estimated ± 1.2 pp, low confidence
71MiMo-V2.6-Flash83.9%estimated ± 1.2 pp, low confidence
72GLM-5V-Turbo83.7%estimated ± 1.2 pp, low confidence
73Gemini 3.1 Flash-Lite83.7%estimated ± 1.1 pp, low confidence
74GPT-5.1-Codex83.5%estimated ± 1.2 pp, low confidence
75GPT-5.1-Codex-Max83.5%estimated ± 1.2 pp, low confidence
76o382.1%estimated ± 1.2 pp, low confidence
77MiMo-V2-Omni82.0%estimated ± 1.2 pp, low confidence
78Gemma 4 26B A4B81.6%estimated ± 1.2 pp, low confidence
79Grok 481.4%estimated ± 1.2 pp, low confidence
80Claude 4.1 Opus Thinking80.8%estimated ± 1.2 pp, low confidence
81Ternary Bonsai 2 27B80.1%measured
82Gemini 2.5 Flash79.3%estimated ± 1.2 pp, low confidence
83GPT-5.4 nano79.3%estimated ± 1.2 pp, low confidence
84Mistral Medium 3.5 128B79.0%estimated ± 1.2 pp, low confidence
85Grok 4.1 Fast (Reasoning)77.9%estimated ± 1.2 pp, low confidence
86Claude 4 Sonnet77.3%estimated ± 1.2 pp, low confidence
87Grok 4.2077.3%estimated ± 1.1 pp, low confidence
88Llama 4 Maverick77.1%estimated ± 1.2 pp, low confidence
89Grok 4 Fast (Reasoning)76.9%estimated ± 1.2 pp, low confidence
90GPT-4.176.5%estimated ± 1.2 pp, low confidence
91Qwen3-Omni-30B-A3B-Thinking75.9%estimated ± 1.2 pp, low confidence
92GPT-4.1 mini74.8%estimated ± 1.2 pp, low confidence
93Claude Opus 4.673.8%estimated ± 1.1 pp, low confidence
94Mistral Small 473.5%estimated ± 1.2 pp, low confidence
95Mistral Small 4 (Reasoning)73.5%estimated ± 1.2 pp, low confidence
96LFM2.5-VL-3B73.1%measured
97Mistral Large 372.7%estimated ± 1.2 pp, low confidence
98Qwen3-Omni-30B-A3B-Instruct72.5%estimated ± 1.2 pp, low confidence
99Gemini 1.5 Pro72.2%estimated ± 1.2 pp, low confidence
100Mistral Medium 370.7%estimated ± 1.2 pp, low confidence
101Llama 4 Scout70.6%estimated ± 1.2 pp, low confidence
102Qwen3.5 397B (Reasoning)70.4%estimated ± 1.2 pp, low confidence
103Gemma 4 E4B69.4%estimated ± 1.2 pp, low confidence
104Grok 4.1 Fast67.0%estimated ± 1.2 pp, low confidence
105Gemma 3 27B66.7%estimated ± 1.2 pp, low confidence
106ZAYA1-VL-8B65.0%measured
107Gemma 4 E2B63.8%estimated ± 1.2 pp, low confidence
108Nova Pro63.5%estimated ± 1.2 pp, low confidence
109Qwen3.5 397B63.2%estimated ± 0.8 pp, low confidence
110North Micro Vision Instruct62.2%measured
111GPT-4o mini61.0%estimated ± 1.2 pp, low confidence
112GPT-4.1 nano59.7%estimated ± 1.2 pp, low confidence
113LFM2.5-VL-450M58.4%measured
114Claude 3 Haiku50.0%estimated ± 1.2 pp, low confidence
115LFM2.5-VL-1.6B-Extract44.9%estimated ± 1.2 pp, low confidence
116Qwen3.6 Plus38.1%estimated ± 0.8 pp, low confidence
117Phi-4 Multimodal Instruct28.0%estimated ± 1.2 pp, low confidence
118dots3-note Preview26.0%estimated ± 0.8 pp, low confidence
119Kimi K2.616.4%estimated ± 0.8 pp, low confidence
120Gemini 3 Pro3.9%estimated ± 0.8 pp, low confidence
121Qwen3.5-122B-A10B1.8%estimated ± 0.8 pp, low confidence
122Qwen3.5-27B1.2%estimated ± 0.8 pp, low confidence
123Qwen3.5-35B-A3B0.0%estimated ± 0.8 pp, low confidence
124GPT-5.20.0%estimated ± 0.8 pp, low confidence
125Claude Opus 4.50.0%estimated ± 0.8 pp, low confidence
126Gemma 4 12B0.0%estimated ± 0.8 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General