benchgap
Vision & documents

AI2D_TEST leaderboard

As of 2026-10-10, the highest measured score on AI2D_TEST is 94.1% by Gemini 3 Pro. 59 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1GPT-5.6 Sol94.7%estimated ± 2.0 pp, low confidence
2Gemini 3 Pro94.1%measured
3Qwen3.5 397B93.9%measured
4GPT-5.593.9%estimated ± 2.0 pp, low confidence
5dots3-note Preview93.7%estimated ± 1.0 pp, medium confidence
6GPT-5.6 Terra93.6%estimated ± 2.0 pp, medium confidence
7Claude Mythos 593.5%estimated ± 0.9 pp, low confidence
8Qwen3.8 Max93.5%estimated ± 0.9 pp, low confidence
9Qwen3.8-Omni-Flash93.5%estimated ± 0.9 pp, low confidence
10Kimi K393.5%estimated ± 0.9 pp, low confidence
11Claude Opus 4.7 (Adaptive)93.5%estimated ± 0.9 pp, low confidence
12Qwen3.8-Flash-Next93.5%estimated ± 0.9 pp, low confidence
13Qwen3.8-27B93.5%estimated ± 0.9 pp, low confidence
14Claude Opus 4.893.5%estimated ± 0.9 pp, low confidence
15GLM-5.3-Flash93.5%estimated ± 0.9 pp, low confidence
16Gemini 3.7 Flash93.5%estimated ± 0.9 pp, low confidence
17Muse Spark 1.193.5%estimated ± 0.9 pp, low confidence
18Claude Sonnet 593.5%estimated ± 0.9 pp, low confidence
19Sakana Fugu-Ultra93.5%estimated ± 0.9 pp, low confidence
20Muse Spark93.5%estimated ± 0.9 pp, low confidence
21Qwen3.7 Plus93.5%estimated ± 0.9 pp, low confidence
22Seed 2.1 Pro93.5%estimated ± 0.9 pp, low confidence
23Sakana Fugu93.5%estimated ± 0.9 pp, low confidence
24Gemini 3.5 Flash93.5%estimated ± 0.9 pp, low confidence
25GPT-5.493.5%estimated ± 0.9 pp, low confidence
26Seed 2.1 Turbo93.5%estimated ± 0.9 pp, low confidence
27Inkling93.5%estimated ± 0.9 pp, medium confidence
28Qwen3.6 Plus93.4%estimated ± 0.9 pp, medium confidence
29Inkling-Small93.4%estimated ± 0.9 pp, medium confidence
30MiMo-V2.593.4%estimated ± 0.9 pp, medium confidence
31Kimi K2.693.4%estimated ± 0.9 pp, medium confidence
32Gemini 3.1 Pro93.4%estimated ± 0.9 pp, medium confidence
33Qwen3.5-27B93.4%estimated ± 0.9 pp, medium confidence
34Muse Glimmer 30B93.2%estimated ± 0.9 pp, medium confidence
35Qwen3.6-27B92.9%estimated ± 0.9 pp, medium confidence
36Qwen3.6-35B-A3B92.7%measured
37GPT-5.6 Luna92.5%estimated ± 2.0 pp, medium confidence
38Grok 4.392.3%estimated ± 2.0 pp, medium confidence
39Pareto 26.992.3%estimated ± 2.0 pp, medium confidence
40GPT-5.292.2%measured
41Claude Opus 4.691.9%estimated ± 2.0 pp, medium confidence
42MiniMax M391.8%estimated ± 1.3 pp, medium confidence
43Gemma 4 31B91.7%estimated ± 2.0 pp, medium confidence
44GPT-5.4 mini91.6%estimated ± 2.0 pp, medium confidence
45Step 5 Preview91.3%estimated ± 2.0 pp, medium confidence
46Kimi K2.5 (Reasoning)91.1%estimated ± 0.9 pp, medium confidence
47Qwen3.5-35B-A3B91.1%estimated ± 0.9 pp, medium confidence
48Claude Sonnet 4.690.8%estimated ± 0.9 pp, medium confidence
49Kimi K2.590.8%measured
50Qwen3.5-122B-A10B90.3%estimated ± 0.9 pp, medium confidence
51Gemma 4 26B A4B90.1%estimated ± 2.0 pp, medium confidence
52Gemma 4 12B90.1%estimated ± 1.0 pp, medium confidence
53Ternary Bonsai 2 27B88.7%estimated ± 1.4 pp, medium confidence
54Interfaze Beta88.6%estimated ± 2.0 pp, medium confidence
55Nemotron 3 Nano Omni 30B A3B88.5%measured
56Gemini 3.1 Flash-Lite87.8%estimated ± 0.9 pp, medium confidence
57Grok 4.2087.8%estimated ± 0.9 pp, low confidence
58Command A+87.8%estimated ± 0.9 pp, low confidence
59Claude Opus 4.587.7%measured
60LFM2.5-VL-3B87.6%estimated ± 1.4 pp, medium confidence
61LFM2.5-VL-450M87.6%estimated ± 1.4 pp, low confidence
62North Micro Vision Instruct87.6%estimated ± 1.4 pp, low confidence
63ZAYA1-VL-8B87.5%measured
64GPT-5.4 nano85.7%estimated ± 2.0 pp, low confidence
65Gemma 4 E4B76.6%estimated ± 2.0 pp, low confidence
66Qwen2.5-VL-32B74.1%estimated ± 2.0 pp, low confidence
67Gemma 4 E2B69.6%estimated ± 2.0 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General