benchgap
Multilingual

PolyMath leaderboard

As of 2026-10-10, the highest measured score on PolyMath is 86.5% by Qwen3.7 Max. 12 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Qwen3.7 Max86.5%measured
2Qwen3.7 Plus84.0%measured
3Claude Opus 4.579.0%measured
4Qwen3.6 Plus73.8%estimated ± 11.4 pp, low confidence
5Qwen3.5 397B73.3%measured
6K-EXAONE 2.071.3%measured
7GLM-550.1%estimated ± 11.4 pp, low confidence
8Nemotron 3 Ultra49.0%estimated ± 11.4 pp, low confidence
9Kimi K2.543.1%measured
10Qwen3.5-122B-A10B42.5%estimated ± 11.4 pp, low confidence
11Qwen3.5-27B42.5%estimated ± 11.4 pp, low confidence
12Qwen3 235B 2507 (Reasoning)39.2%estimated ± 11.4 pp, low confidence
13Qwen3.5-35B-A3B39.2%estimated ± 11.4 pp, low confidence
14Qwen3 235B 250738.4%estimated ± 11.4 pp, low confidence
15GPT-4.138.2%estimated ± 11.4 pp, low confidence
16DeepSeek V3 032438.2%estimated ± 11.4 pp, low confidence
17GPT-4o38.2%estimated ± 11.4 pp, low confidence
18Phi-438.2%estimated ± 11.4 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General