benchgap
Vision & documents

CharXiv w/o tools leaderboard

As of 2026-10-07, the highest measured score on CharXiv w/o tools is 88.9% by Claude Mythos 5. 108 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Mythos 588.9%measured
2Gemini 3.1 Pro88.6%estimated ± 1.8 pp, low confidence
3Qwen3.8 Max88.4%measured
4Gemini 3.5 Flash88.3%estimated ± 1.8 pp, low confidence
5GPT-5.6 Sol87.7%estimated ± 1.8 pp, low confidence
6Claude Opus 5.587.2%estimated ± 2.4 pp, medium confidence
7GPT-6 Astra86.7%estimated ± 2.4 pp, medium confidence
8GPT-6.1 Sol86.2%estimated ± 2.4 pp, medium confidence
9Gemini 3.8 Flash86.2%measured
10GPT-5.485.7%estimated ± 1.8 pp, medium confidence
11GPT-5.585.7%estimated ± 1.8 pp, medium confidence
12Claude Opus 585.5%estimated ± 2.4 pp, high confidence
13GPT-5.6 Terra85.2%estimated ± 1.8 pp, medium confidence
14Muse Spark84.9%estimated ± 1.8 pp, medium confidence
15Kimi K384.8%measured
16Qwen3.8-Flash-Next84.6%measured
17Gemini 3.6 Flash84.6%estimated ± 2.4 pp, high confidence
18Gemini 3.7 Flash84.5%measured
19GPT-6 Sol84.4%estimated ± 2.4 pp, high confidence
20Qwen3.8 Max Preview84.4%estimated ± 2.4 pp, high confidence
21Seed 2.1 Pro83.8%estimated ± 1.7 pp, medium confidence
22Qwen3.7 Plus83.7%estimated ± 1.7 pp, medium confidence
23Seed 2.1 Turbo83.7%estimated ± 1.7 pp, medium confidence
24Qwen3.5 397B83.7%estimated ± 1.7 pp, medium confidence
25Qwen3.6 Plus83.7%estimated ± 1.7 pp, medium confidence
26Claude Opus 4.583.7%estimated ± 1.7 pp, low confidence
27Gemini 3 Pro83.7%estimated ± 1.7 pp, low confidence
28Gemma 4 12B83.7%estimated ± 1.7 pp, low confidence
29GPT-5.283.7%estimated ± 1.7 pp, low confidence
30Kimi K2.683.7%estimated ± 1.7 pp, low confidence
31Qwen3.5-122B-A10B83.7%estimated ± 1.7 pp, low confidence
32Qwen3.5-27B83.7%estimated ± 1.7 pp, low confidence
33Qwen3.5-35B-A3B83.7%estimated ± 1.7 pp, low confidence
34Qwen3.8-27B83.7%measured
35Qwen3.8-Omni-Flash83.5%measured
36dots3-note Preview83.1%measured
37Grok 4.582.9%estimated ± 2.4 pp, high confidence
38Kimi K2.582.8%estimated ± 1.8 pp, medium confidence
39Kimi K2.5 (Reasoning)82.8%estimated ± 1.8 pp, medium confidence
40GPT-5.6 Luna82.7%estimated ± 1.8 pp, medium confidence
41GPT-6 Luna82.5%estimated ± 2.4 pp, high confidence
42Grok 4.382.4%estimated ± 1.8 pp, medium confidence
43MiniMax M382.4%estimated ± 1.8 pp, medium confidence
44Pareto 26.982.3%estimated ± 1.8 pp, medium confidence
45Apodex 1.182.2%estimated ± 2.4 pp, high confidence
46Apodex 1.1 Mini82.2%estimated ± 2.4 pp, high confidence
47MiMo-V2.582.1%estimated ± 1.8 pp, medium confidence
48Claude Opus 4.7 (Adaptive)82.1%measured
49Gemini 3.5 Flash-Lite82.1%estimated ± 2.4 pp, high confidence
50Ling 3.0 Flash VL82.1%estimated ± 2.4 pp, high confidence
51Gemini 3 Flash81.8%estimated ± 2.4 pp, high confidence
52GPT-5.3 Codex81.8%estimated ± 2.4 pp, high confidence
53GLM-5.3-Flash81.5%estimated ± 2.3 pp, high confidence
54Claude Opus 4.681.5%estimated ± 1.8 pp, medium confidence
55Gemma 4 31B81.1%estimated ± 1.8 pp, medium confidence
56DeepSeek V4.1 Flash80.8%estimated ± 2.4 pp, high confidence
57GPT-5.4 mini80.7%estimated ± 1.8 pp, medium confidence
58Muse Spark 1.180.5%estimated ± 2.3 pp, high confidence
59Claude Opus 4.880.5%measured
60Claude Opus 4.780.5%estimated ± 2.4 pp, high confidence
61Mistral Large 480.5%estimated ± 2.4 pp, high confidence
62GPT-5.2-Codex80.4%estimated ± 2.4 pp, high confidence
63Step 5 Preview80.1%estimated ± 1.8 pp, medium confidence
64GPT-5.179.9%estimated ± 2.4 pp, high confidence
65Qwen3.6-27B79.9%estimated ± 1.8 pp, medium confidence
66Claude Opus 4.6 (Adaptive)79.8%estimated ± 2.4 pp, high confidence
67Step 3.7 Flash79.8%estimated ± 2.4 pp, high confidence
68Gemini 2.5 Pro79.5%estimated ± 2.4 pp, high confidence
69Qwen3.6-35B-A3B79.3%estimated ± 1.8 pp, medium confidence
70Sakana Fugu-Ultra79.2%estimated ± 2.3 pp, high confidence
71Grok 4.2079.2%estimated ± 1.8 pp, medium confidence
72GPT-5 (medium)79.1%estimated ± 2.4 pp, high confidence
73GPT-5 (high)79.1%estimated ± 2.4 pp, high confidence
74Claude Opus 4.5 Thinking78.9%estimated ± 2.4 pp, high confidence
75Sakana Fugu78.5%estimated ± 2.3 pp, high confidence
76MiMo-V2.6-Flash78.4%estimated ± 2.4 pp, medium confidence
77GLM-5V-Turbo78.2%estimated ± 2.4 pp, medium confidence
78Inkling78.1%measured
79GPT-5.1-Codex78.0%estimated ± 2.4 pp, medium confidence
80GPT-5.1-Codex-Max78.0%estimated ± 2.4 pp, medium confidence
81Muse Glimmer 30B77.9%estimated ± 1.8 pp, medium confidence
82Gemma 4 26B A4B77.7%estimated ± 1.8 pp, medium confidence
83Inkling-Small77.4%measured
84Claude Sonnet 4.677.3%estimated ± 2.3 pp, medium confidence
85Nemotron 3 Nano Omni 30B A3B77.3%estimated ± 2.3 pp, medium confidence
86Gemini 3.1 Flash-Lite77.2%estimated ± 2.3 pp, medium confidence
87Claude Sonnet 577.0%measured
88o376.4%estimated ± 2.4 pp, medium confidence
89MiMo-V2-Omni76.2%estimated ± 2.4 pp, medium confidence
90Grok 475.5%estimated ± 2.4 pp, medium confidence
91Claude 4.1 Opus Thinking74.9%estimated ± 2.4 pp, medium confidence
92Interfaze Beta74.8%estimated ± 1.8 pp, low confidence
93Gemini 2.5 Flash73.2%estimated ± 2.4 pp, medium confidence
94Mistral Medium 3.5 128B72.8%estimated ± 2.4 pp, medium confidence
95Grok 4.1 Fast (Reasoning)71.6%estimated ± 2.4 pp, medium confidence
96Claude 4 Sonnet71.0%estimated ± 2.4 pp, medium confidence
97Llama 4 Maverick70.7%estimated ± 2.4 pp, medium confidence
98Grok 4 Fast (Reasoning)70.5%estimated ± 2.4 pp, medium confidence
99GPT-4.170.1%estimated ± 2.4 pp, medium confidence
100GPT-5.4 nano69.4%estimated ± 1.8 pp, low confidence
101Qwen3-Omni-30B-A3B-Thinking69.3%estimated ± 2.4 pp, medium confidence
102GPT-4.1 mini68.2%estimated ± 2.4 pp, medium confidence
103Mistral Small 466.7%estimated ± 2.4 pp, medium confidence
104Mistral Small 4 (Reasoning)66.7%estimated ± 2.4 pp, medium confidence
105Command A+66.0%estimated ± 1.8 pp, low confidence
106Mistral Large 365.8%estimated ± 2.4 pp, medium confidence
107Qwen3-Omni-30B-A3B-Instruct65.7%estimated ± 2.4 pp, medium confidence
108Gemini 1.5 Pro65.3%estimated ± 2.4 pp, medium confidence
109Mistral Medium 363.7%estimated ± 2.4 pp, medium confidence
110Llama 4 Scout63.6%estimated ± 2.4 pp, medium confidence
111Qwen3.5 397B (Reasoning)63.4%estimated ± 2.4 pp, medium confidence
112Gemma 4 E4B62.3%estimated ± 2.4 pp, medium confidence
113Grok 4.1 Fast59.8%estimated ± 2.4 pp, medium confidence
114Gemma 3 27B59.4%estimated ± 2.4 pp, medium confidence
115Gemma 4 E2B56.4%estimated ± 2.4 pp, medium confidence
116Nova Pro56.1%estimated ± 2.4 pp, medium confidence
117GPT-4o mini53.5%estimated ± 2.4 pp, medium confidence
118GPT-4.1 nano52.2%estimated ± 2.4 pp, medium confidence
119Claude 3 Haiku42.7%estimated ± 2.4 pp, medium confidence
120LFM2.5-VL-1.6B-Extract37.9%estimated ± 2.4 pp, medium confidence
121LFM2.5-VL-3B30.8%estimated ± 1.8 pp, low confidence
122Phi-4 Multimodal Instruct22.7%estimated ± 2.4 pp, medium confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General