benchgap
Vision & documents

We-Math leaderboard

As of 2026-10-10, the highest measured score on We-Math is 87.9% by Qwen3.5 397B. 62 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Qwen3.8 Max97.3%estimated ± 1.9 pp, low confidence
2Kimi K396.1%estimated ± 1.9 pp, low confidence
3Seed 2.1 Pro93.9%estimated ± 1.9 pp, low confidence
4Qwen3.8-Omni-Flash92.8%estimated ± 1.9 pp, low confidence
5Qwen3.8-Flash-Next91.2%estimated ± 1.9 pp, low confidence
6Qwen3.7 Plus90.8%estimated ± 1.9 pp, low confidence
7Seed 2.1 Turbo90.6%estimated ± 1.9 pp, low confidence
8Qwen3.8-27B90.5%estimated ± 1.9 pp, low confidence
9Qwen3.5 397B87.9%measured
10Qwen3.6 Plus87.8%estimated ± 1.9 pp, medium confidence
11dots3-note Preview87.4%estimated ± 1.9 pp, medium confidence
12Kimi K2.687.0%estimated ± 1.9 pp, medium confidence
13Gemini 3 Pro86.9%measured
14Gemini 3.1 Pro85.5%estimated ± 6.2 pp, low confidence
15Qwen3.5-122B-A10B85.5%estimated ± 1.9 pp, medium confidence
16Gemini 3.5 Flash85.3%estimated ± 6.2 pp, low confidence
17MiMo-V2.585.3%estimated ± 3.7 pp, medium confidence
18Qwen3.6-27B85.3%estimated ± 3.7 pp, medium confidence
19Qwen3.5-27B85.2%estimated ± 1.9 pp, medium confidence
20GPT-5.6 Sol85.0%estimated ± 6.2 pp, low confidence
21Kimi K2.584.7%measured
22Claude Mythos 584.6%estimated ± 7.9 pp, low confidence
23Claude Opus 4.7 (Adaptive)84.6%estimated ± 7.9 pp, low confidence
24Claude Opus 4.884.6%estimated ± 7.9 pp, low confidence
25GLM-5.3-Flash84.6%estimated ± 7.9 pp, low confidence
26Gemini 3.7 Flash84.6%estimated ± 7.9 pp, low confidence
27Muse Spark 1.184.6%estimated ± 7.9 pp, low confidence
28Claude Sonnet 584.6%estimated ± 7.9 pp, low confidence
29Sakana Fugu-Ultra84.6%estimated ± 7.9 pp, low confidence
30Sakana Fugu84.6%estimated ± 7.9 pp, low confidence
31Claude Sonnet 4.684.6%estimated ± 7.9 pp, low confidence
32Gemini 3.1 Flash-Lite84.5%estimated ± 7.9 pp, low confidence
33GPT-5.483.9%estimated ± 6.2 pp, low confidence
34GPT-5.583.9%estimated ± 6.2 pp, low confidence
35Qwen3.6-35B-A3B83.7%estimated ± 3.7 pp, medium confidence
36GPT-5.6 Terra83.6%estimated ± 6.2 pp, low confidence
37Muse Spark83.4%estimated ± 6.2 pp, low confidence
38Qwen3.5-35B-A3B82.4%estimated ± 1.9 pp, medium confidence
39Kimi K2.5 (Reasoning)82.3%estimated ± 6.2 pp, low confidence
40GPT-5.6 Luna82.2%estimated ± 6.2 pp, low confidence
41Grok 4.382.0%estimated ± 6.2 pp, low confidence
42Pareto 26.982.0%estimated ± 6.2 pp, low confidence
43MiniMax M381.9%estimated ± 3.7 pp, medium confidence
44Claude Opus 4.681.5%estimated ± 6.2 pp, low confidence
45Gemma 4 31B81.3%estimated ± 6.2 pp, low confidence
46GPT-5.4 mini81.1%estimated ± 6.2 pp, low confidence
47Step 5 Preview80.7%estimated ± 6.2 pp, low confidence
48Grok 4.2080.2%estimated ± 6.2 pp, low confidence
49Inkling-Small79.4%estimated ± 6.2 pp, low confidence
50Muse Glimmer 30B79.4%estimated ± 6.2 pp, low confidence
51Gemma 4 26B A4B79.3%estimated ± 6.2 pp, low confidence
52Inkling79.1%estimated ± 6.2 pp, low confidence
53GPT-5.279.0%measured
54Ternary Bonsai 2 27B78.2%estimated ± 5.4 pp, low confidence
55Interfaze Beta77.5%estimated ± 6.2 pp, low confidence
56Gemma 4 12B76.9%estimated ± 1.9 pp, medium confidence
57Nemotron 3 Nano Omni 30B A3B75.9%estimated ± 5.1 pp, low confidence
58GPT-5.4 nano74.1%estimated ± 6.2 pp, low confidence
59Command A+71.9%estimated ± 6.2 pp, low confidence
60ZAYA1-VL-8B70.7%estimated ± 4.0 pp, low confidence
61Claude Opus 4.570.0%measured
62LFM2.5-VL-3B69.1%estimated ± 4.0 pp, low confidence
63Gemma 4 E4B63.8%estimated ± 6.2 pp, low confidence
64Qwen2.5-VL-32B61.2%estimated ± 6.2 pp, low confidence
65Gemma 4 E2B56.5%estimated ± 6.2 pp, low confidence
66North Micro Vision Instruct43.6%estimated ± 5.4 pp, low confidence
67LFM2.5-VL-450M38.7%estimated ± 4.0 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General