benchgap
Vision & documents

CharXiv leaderboard

As of 2026-10-07, the highest measured score on CharXiv is 93.5% by Claude Mythos 5. 91 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Mythos 593.5%measured
2Qwen3.8 Max93.5%measured
3Qwen3.8-Omni-Flash91.4%measured
4Gemini 3.8 Flash91.4%estimated ± 2.5 pp, high confidence
5Kimi K391.3%measured
6Claude Opus 4.7 (Adaptive)91.0%measured
7Qwen3.8-Flash-Next90.6%measured
8Qwen3.8-27B90.2%measured
9Claude Opus 4.889.9%measured
10dots3-note Preview89.5%estimated ± 2.5 pp, high confidence
11GLM-5.3-Flash89.4%measured
12Gemini 3.7 Flash88.7%measured
13GPT-6 Astra88.5%estimated ± 3.9 pp, medium confidence
14Muse Spark 1.188.4%measured
15Claude Sonnet 588.3%measured
16Sakana Fugu-Ultra86.6%measured
17Muse Spark86.4%measured
18GPT-5.6 Sol86.2%estimated ± 5.4 pp, medium confidence
19Claude Opus 5.586.0%estimated ± 7.5 pp, low confidence
20GPT-6.1 Sol86.0%estimated ± 7.5 pp, low confidence
21Claude Opus 586.0%estimated ± 7.5 pp, low confidence
22Gemini 3.6 Flash86.0%estimated ± 7.5 pp, low confidence
23GPT-6 Sol86.0%estimated ± 7.5 pp, low confidence
24Qwen3.8 Max Preview86.0%estimated ± 7.5 pp, low confidence
25Qwen3.7 Plus85.9%measured
26Grok 4.585.8%estimated ± 7.5 pp, low confidence
27GPT-6 Luna85.7%estimated ± 7.5 pp, low confidence
28Claude Opus 4.685.7%estimated ± 3.9 pp, high confidence
29Apodex 1.185.6%estimated ± 7.5 pp, low confidence
30Apodex 1.1 Mini85.6%estimated ± 7.5 pp, low confidence
31Gemini 3.5 Flash-Lite85.6%estimated ± 7.5 pp, low confidence
32Ling 3.0 Flash VL85.6%estimated ± 7.5 pp, low confidence
33Gemini 3 Flash85.4%estimated ± 7.5 pp, low confidence
34Seed 2.1 Pro85.4%measured
35GPT-5.3 Codex85.4%estimated ± 7.5 pp, low confidence
36Sakana Fugu85.1%measured
37GPT-5.584.5%estimated ± 5.4 pp, medium confidence
38DeepSeek V4.1 Flash84.4%estimated ± 7.5 pp, low confidence
39Gemini 3.5 Flash84.2%measured
40GPT-5.6 Terra84.0%estimated ± 5.4 pp, medium confidence
41Claude Opus 4.783.8%estimated ± 7.5 pp, low confidence
42Mistral Large 483.8%estimated ± 7.5 pp, low confidence
43GPT-5.2-Codex83.7%estimated ± 7.5 pp, low confidence
44GPT-5.482.8%measured
45Seed 2.1 Turbo82.5%measured
46GPT-5.182.4%estimated ± 7.5 pp, low confidence
47Claude Opus 4.6 (Adaptive)82.3%estimated ± 7.5 pp, low confidence
48GPT-5.282.1%measured
49Inkling82.0%measured
50Kimi K2.581.6%estimated ± 5.4 pp, medium confidence
51Kimi K2.5 (Reasoning)81.6%estimated ± 5.4 pp, medium confidence
52Qwen3.6 Plus81.5%measured
53GPT-5.6 Luna81.5%estimated ± 5.4 pp, medium confidence
54Gemini 3 Pro81.4%measured
55Holo2-235B-A22B81.4%estimated ± 3.9 pp, high confidence
56Inkling-Small81.3%measured
57Gemini 2.5 Pro81.3%estimated ± 7.5 pp, low confidence
58Qwen3.5-27B81.2%estimated ± 3.3 pp, high confidence
59Grok 4.381.1%estimated ± 5.4 pp, medium confidence
60MiMo-V2.581.0%measured
61Pareto 26.981.0%estimated ± 5.4 pp, medium confidence
62Qwen3.5 397B80.8%measured
63Kimi K2.680.4%measured
64Step 3.7 Flash80.3%estimated ± 4.9 pp, high confidence
65Gemini 3.1 Pro80.2%measured
66GPT-5 (medium)79.9%estimated ± 7.5 pp, low confidence
67GPT-5 (high)79.7%estimated ± 7.5 pp, low confidence
68Gemma 4 31B79.6%estimated ± 5.4 pp, medium confidence
69Holo2-30B-A3B79.6%estimated ± 3.9 pp, high confidence
70GPT-5.4 mini79.2%estimated ± 5.4 pp, medium confidence
71Claude Opus 4.5 Thinking79.2%estimated ± 7.5 pp, low confidence
72Muse Glimmer 30B78.8%measured
73Qwen3.5-35B-A3B78.7%estimated ± 3.3 pp, high confidence
74Step 5 Preview78.4%estimated ± 5.4 pp, medium confidence
75Qwen3.6-27B78.4%measured
76Qwen3.6-35B-A3B78.0%measured
77Claude Sonnet 4.677.4%measured
78Qwen3.5-122B-A10B77.2%measured
79MiMo-V2.6-Flash76.9%estimated ± 7.5 pp, low confidence
80Holo2-8B76.3%estimated ± 3.9 pp, high confidence
81Nemotron 3 Nano Omni 30B A3B76.3%measured
82GLM-5V-Turbo76.2%estimated ± 7.5 pp, low confidence
83GPT-5.1-Codex75.5%estimated ± 7.5 pp, low confidence
84GPT-5.1-Codex-Max75.5%estimated ± 7.5 pp, low confidence
85Holo2-4B75.5%estimated ± 3.9 pp, high confidence
86Gemma 4 26B A4B75.3%estimated ± 5.4 pp, medium confidence
87Gemma 4 12B73.9%estimated ± 3.3 pp, high confidence
88Gemini 3.1 Flash-Lite73.2%measured
89o371.9%estimated ± 7.5 pp, low confidence
90MiniMax M371.8%estimated ± 2.7 pp, low confidence
91MiMo-V2-Omni71.7%estimated ± 7.5 pp, low confidence
92Grok 471.0%estimated ± 7.5 pp, low confidence
93Interfaze Beta70.8%estimated ± 5.4 pp, medium confidence
94Claude 4.1 Opus Thinking70.7%estimated ± 7.5 pp, low confidence
95Gemini 2.5 Flash70.4%estimated ± 7.5 pp, low confidence
96Mistral Medium 3.5 128B70.4%estimated ± 7.5 pp, low confidence
97Grok 4.1 Fast (Reasoning)70.4%estimated ± 7.5 pp, low confidence
98Claude 4 Sonnet70.4%estimated ± 7.5 pp, low confidence
99Llama 4 Maverick70.4%estimated ± 7.5 pp, low confidence
100Grok 4 Fast (Reasoning)70.4%estimated ± 7.5 pp, low confidence
101GPT-4.170.4%estimated ± 7.5 pp, low confidence
102Qwen3-Omni-30B-A3B-Thinking70.3%estimated ± 7.5 pp, low confidence
103GPT-4.1 mini70.3%estimated ± 7.5 pp, low confidence
104Mistral Small 470.3%estimated ± 7.5 pp, low confidence
105Mistral Small 4 (Reasoning)70.3%estimated ± 7.5 pp, low confidence
106Gemini 1.5 Pro70.3%estimated ± 7.5 pp, low confidence
107Mistral Large 370.3%estimated ± 7.5 pp, low confidence
108Qwen3-Omni-30B-A3B-Instruct70.3%estimated ± 7.5 pp, low confidence
109Claude 3 Haiku70.3%estimated ± 7.5 pp, low confidence
110Gemma 3 27B70.3%estimated ± 7.5 pp, low confidence
111Gemma 4 E2B70.3%estimated ± 7.5 pp, low confidence
112Gemma 4 E4B70.3%estimated ± 7.5 pp, low confidence
113GPT-4.1 nano70.3%estimated ± 7.5 pp, low confidence
114GPT-4o mini70.3%estimated ± 7.5 pp, low confidence
115Grok 4.1 Fast70.3%estimated ± 7.5 pp, low confidence
116LFM2.5-VL-1.6B-Extract70.3%estimated ± 7.5 pp, low confidence
117Llama 4 Scout70.3%estimated ± 7.5 pp, low confidence
118Mistral Medium 370.3%estimated ± 7.5 pp, low confidence
119Nova Pro70.3%estimated ± 7.5 pp, low confidence
120Phi-4 Multimodal Instruct70.3%estimated ± 7.5 pp, low confidence
121Qwen3.5 397B (Reasoning)70.3%estimated ± 7.5 pp, low confidence
122Claude Opus 4.568.5%measured
123Ternary Bonsai 2 27B65.0%estimated ± 4.1 pp, low confidence
124GPT-5.4 nano61.0%estimated ± 5.4 pp, medium confidence
125Grok 4.2060.9%measured
126Command A+52.7%measured
127LFM2.5-VL-3B41.4%estimated ± 4.1 pp, low confidence
128ZAYA1-VL-8B14.1%estimated ± 4.1 pp, low confidence
129North Micro Vision Instruct4.7%estimated ± 4.1 pp, low confidence
130LFM2.5-VL-450M0.0%estimated ± 4.1 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General