benchgap
Vision & documents

SimpleVQA leaderboard

As of 2026-10-07, the highest measured score on SimpleVQA is 81.7% by Qwen3.7 Plus. 57 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 4.599.0%estimated ± 5.7 pp, low confidence
2Holo2-4B96.2%estimated ± 5.7 pp, low confidence
3Nemotron 3 Nano Omni 30B A3B95.9%estimated ± 5.7 pp, low confidence
4Holo2-8B95.5%estimated ± 5.7 pp, low confidence
5Qwen3.5 397B91.7%estimated ± 5.7 pp, low confidence
6Holo2-30B-A3B91.3%estimated ± 5.7 pp, low confidence
7Qwen3.6 Plus89.7%estimated ± 5.7 pp, low confidence
8Holo2-235B-A22B87.7%estimated ± 5.7 pp, low confidence
9Gemini 3 Pro85.6%estimated ± 5.7 pp, low confidence
10Muse Glimmer 30B82.7%estimated ± 5.7 pp, low confidence
11Qwen3.7 Plus81.7%measured
12Step 3.7 Flash79.2%measured
13Qwen3.8 Max75.0%measured
14Seed 2.1 Turbo74.4%estimated ± 2.5 pp, low confidence
15Seed 2.1 Pro74.3%estimated ± 2.5 pp, low confidence
16Kimi K373.9%estimated ± 2.5 pp, medium confidence
17Claude Mythos 573.6%estimated ± 7.8 pp, low confidence
18Claude Opus 4.672.8%estimated ± 5.7 pp, low confidence
19dots3-note Preview72.5%measured
20Gemini 3.1 Pro72.4%measured
21Claude Opus 4.7 (Adaptive)72.4%estimated ± 7.8 pp, low confidence
22Qwen3.8-Flash-Next72.1%estimated ± 7.4 pp, low confidence
23GLM-5.3-Flash71.5%estimated ± 7.8 pp, low confidence
24Muse Spark71.3%measured
25Qwen3.8-Omni-Flash71.2%estimated ± 7.4 pp, low confidence
26Gemini 3.7 Flash71.2%estimated ± 7.8 pp, low confidence
27Muse Spark 1.171.0%estimated ± 7.8 pp, low confidence
28Claude Sonnet 571.0%estimated ± 7.8 pp, low confidence
29Sakana Fugu-Ultra70.1%estimated ± 7.8 pp, low confidence
30GPT-5.6 Sol69.8%estimated ± 8.2 pp, medium confidence
31Sakana Fugu69.3%estimated ± 7.8 pp, low confidence
32Gemini 3.5 Flash68.8%estimated ± 7.8 pp, low confidence
33GPT-5.568.8%estimated ± 8.2 pp, medium confidence
34GPT-5.6 Terra68.5%estimated ± 8.2 pp, medium confidence
35GPT-5.267.7%estimated ± 7.8 pp, low confidence
36Inkling67.6%estimated ± 7.8 pp, low confidence
37Qwen3.8-27B67.6%estimated ± 7.4 pp, low confidence
38Kimi K2.567.3%estimated ± 8.2 pp, medium confidence
39Kimi K2.5 (Reasoning)67.3%estimated ± 8.2 pp, medium confidence
40Inkling-Small67.2%estimated ± 7.8 pp, low confidence
41GPT-5.6 Luna67.2%estimated ± 8.2 pp, medium confidence
42MiMo-V2.567.1%estimated ± 7.8 pp, low confidence
43Grok 4.367.0%estimated ± 8.2 pp, medium confidence
44MiniMax M367.0%estimated ± 8.2 pp, medium confidence
45Pareto 26.967.0%estimated ± 8.2 pp, medium confidence
46Kimi K2.666.7%estimated ± 7.8 pp, low confidence
47Gemma 4 31B66.3%estimated ± 8.2 pp, medium confidence
48GPT-5.4 mini66.2%estimated ± 8.2 pp, medium confidence
49Step 5 Preview65.8%estimated ± 8.2 pp, medium confidence
50Claude Opus 4.865.6%estimated ± 5.7 pp, low confidence
51Claude Sonnet 4.665.1%estimated ± 7.8 pp, low confidence
52Qwen3.5-122B-A10B65.0%estimated ± 7.8 pp, low confidence
53Gemma 4 26B A4B64.5%estimated ± 8.2 pp, medium confidence
54Interfaze Beta62.9%estimated ± 8.2 pp, medium confidence
55Gemini 3.1 Flash-Lite62.6%estimated ± 7.8 pp, low confidence
56Gemma 4 12B61.7%estimated ± 8.2 pp, medium confidence
57GPT-5.461.1%measured
58GPT-5.4 nano59.8%estimated ± 8.2 pp, medium confidence
59Qwen3.6-35B-A3B58.9%measured
60GPT-6 Astra58.1%estimated ± 5.7 pp, low confidence
61Grok 4.2057.4%measured
62Qwen3.6-27B56.1%measured
63Ternary Bonsai 2 27B54.3%estimated ± 9.5 pp, low confidence
64Command A+49.4%estimated ± 7.8 pp, low confidence
65LFM2.5-VL-3B35.4%measured
66ZAYA1-VL-8B22.9%estimated ± 9.5 pp, low confidence
67North Micro Vision Instruct18.4%estimated ± 9.5 pp, low confidence
68LFM2.5-VL-450M13.3%estimated ± 9.5 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General