benchgap
Vision & documents

DynaMath leaderboard

As of 2026-10-10, the highest measured score on DynaMath is 86.8% by GPT-5.2. 62 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Mythos 588.5%estimated ± 0.8 pp, low confidence
2Qwen3.8 Max88.5%estimated ± 0.8 pp, low confidence
3Qwen3.8-Omni-Flash88.2%estimated ± 0.8 pp, low confidence
4Kimi K388.2%estimated ± 0.8 pp, low confidence
5Claude Opus 4.7 (Adaptive)88.2%estimated ± 0.8 pp, low confidence
6Qwen3.8-Flash-Next88.1%estimated ± 0.8 pp, low confidence
7Qwen3.8-27B88.0%estimated ± 0.8 pp, low confidence
8Claude Opus 4.888.0%estimated ± 0.8 pp, low confidence
9GLM-5.3-Flash87.9%estimated ± 0.8 pp, low confidence
10Gemini 3.7 Flash87.8%estimated ± 0.8 pp, low confidence
11Muse Spark 1.187.7%estimated ± 0.8 pp, low confidence
12Claude Sonnet 587.7%estimated ± 0.8 pp, low confidence
13Sakana Fugu-Ultra87.4%estimated ± 0.8 pp, low confidence
14Muse Spark87.4%estimated ± 0.8 pp, low confidence
15Qwen3.7 Plus87.2%estimated ± 0.8 pp, low confidence
16GPT-5.6 Sol87.2%estimated ± 2.7 pp, low confidence
17Seed 2.1 Pro87.1%estimated ± 0.8 pp, low confidence
18Sakana Fugu87.1%estimated ± 0.8 pp, low confidence
19Gemini 3.5 Flash86.9%estimated ± 0.8 pp, low confidence
20GPT-5.286.8%measured
21GPT-5.586.6%estimated ± 2.7 pp, low confidence
22GPT-5.486.5%estimated ± 0.8 pp, low confidence
23Seed 2.1 Turbo86.4%estimated ± 0.8 pp, low confidence
24GPT-5.6 Terra86.4%estimated ± 2.7 pp, medium confidence
25Qwen3.5 397B86.3%measured
26Inkling86.3%estimated ± 0.8 pp, medium confidence
27Qwen3.6 Plus86.1%estimated ± 0.8 pp, medium confidence
28Inkling-Small86.1%estimated ± 0.8 pp, medium confidence
29MiMo-V2.586.0%estimated ± 0.8 pp, medium confidence
30Step 3.7 Flash86.0%estimated ± 3.3 pp, medium confidence
31Kimi K2.685.8%estimated ± 0.8 pp, medium confidence
32Gemini 3.1 Pro85.7%estimated ± 0.8 pp, medium confidence
33dots3-note Preview85.7%estimated ± 2.7 pp, medium confidence
34Qwen3.6-27B85.6%measured
35Qwen3.5-27B85.5%estimated ± 0.8 pp, medium confidence
36GPT-5.6 Luna85.4%estimated ± 2.7 pp, medium confidence
37Grok 4.385.3%estimated ± 2.7 pp, medium confidence
38MiniMax M385.3%estimated ± 2.7 pp, medium confidence
39Muse Glimmer 30B85.3%estimated ± 0.8 pp, medium confidence
40Pareto 26.985.2%estimated ± 2.7 pp, medium confidence
41Gemini 3 Pro85.1%measured
42Qwen3.6-35B-A3B85.0%estimated ± 0.8 pp, medium confidence
43Claude Opus 4.684.9%estimated ± 2.7 pp, medium confidence
44Kimi K2.5 (Reasoning)84.8%estimated ± 0.8 pp, medium confidence
45Qwen3.5-35B-A3B84.8%estimated ± 0.8 pp, medium confidence
46Claude Sonnet 4.684.7%estimated ± 0.8 pp, medium confidence
47Gemma 4 31B84.7%estimated ± 2.7 pp, medium confidence
48Qwen3.5-122B-A10B84.6%estimated ± 0.8 pp, medium confidence
49GPT-5.4 mini84.5%estimated ± 2.7 pp, medium confidence
50Kimi K2.584.4%measured
51Nemotron 3 Nano Omni 30B A3B84.2%estimated ± 0.8 pp, medium confidence
52Step 5 Preview84.2%estimated ± 2.7 pp, medium confidence
53Ternary Bonsai 2 27B83.1%estimated ± 1.2 pp, medium confidence
54Gemma 4 26B A4B82.8%estimated ± 2.7 pp, medium confidence
55Gemini 3.1 Flash-Lite82.7%estimated ± 0.8 pp, medium confidence
56Interfaze Beta80.8%estimated ± 2.7 pp, medium confidence
57Claude Opus 4.579.7%measured
58Gemma 4 12B79.0%estimated ± 2.7 pp, low confidence
59GPT-5.4 nano75.7%estimated ± 2.7 pp, low confidence
60LFM2.5-VL-3B75.0%estimated ± 1.2 pp, low confidence
61Grok 4.2072.3%estimated ± 0.8 pp, low confidence
62Command A+60.7%estimated ± 0.8 pp, low confidence
63ZAYA1-VL-8B60.6%estimated ± 1.2 pp, low confidence
64North Micro Vision Instruct54.4%estimated ± 1.2 pp, low confidence
65Gemma 4 E4B49.8%estimated ± 2.7 pp, low confidence
66LFM2.5-VL-450M45.3%estimated ± 1.2 pp, low confidence
67Qwen2.5-VL-32B41.4%estimated ± 2.7 pp, low confidence
68Gemma 4 E2B27.0%estimated ± 2.7 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General