benchgap
Math

FrontierMath (legacy) leaderboard

As of 2026-10-07, the highest measured score on FrontierMath (legacy) is 89.0% by GPT-5.6 Sol. 46 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1GPT-6 Astra90.7%estimated ± 2.3 pp, low confidence
2GPT-5.6 Sol89.0%measured
3GPT-5.6 Terra84.9%measured
4GPT-5.6 Luna78.6%measured
5GPT-5.5 Pro52.4%measured
6GPT-5.551.7%measured
7GPT-5.4 Pro50.0%measured
8GPT-5.448.1%estimated ± 0.7 pp, low confidence
9Claude Opus 4.847.7%estimated ± 0.7 pp, low confidence
10Claude Opus 4.744.3%estimated ± 0.7 pp, low confidence
11Claude Opus 4.7 (Adaptive)43.8%measured
12Claude Opus 4.641.3%estimated ± 0.7 pp, low confidence
13GPT-5.241.3%estimated ± 0.7 pp, low confidence
14Muse Spark39.6%estimated ± 0.7 pp, low confidence
15Gemini 3.5 Flash39.6%estimated ± 0.7 pp, low confidence
16Kimi K2.639.6%estimated ± 0.7 pp, low confidence
17Gemini 3 Pro38.2%estimated ± 0.7 pp, low confidence
18Gemini 3.1 Pro37.5%estimated ± 0.7 pp, low confidence
19Gemini 3 Flash36.3%estimated ± 0.7 pp, low confidence
20GLM-5.134.1%estimated ± 0.7 pp, low confidence
21Claude Sonnet 4.633.1%estimated ± 0.7 pp, low confidence
22GPT-5.131.8%estimated ± 0.7 pp, low confidence
23GPT-5.4 mini29.0%estimated ± 0.7 pp, low confidence
24Kimi K2.528.7%estimated ± 0.7 pp, low confidence
25Qwen3.6 Plus27.0%estimated ± 0.7 pp, low confidence
26GPT-5.4 nano26.7%estimated ± 0.7 pp, low confidence
27o4-mini (high)25.6%estimated ± 0.7 pp, low confidence
28Qwen 3.6 Max (preview)23.9%estimated ± 0.7 pp, low confidence
29DeepSeek V3.222.9%estimated ± 0.7 pp, low confidence
30Kimi K222.3%estimated ± 0.7 pp, low confidence
31Qwen3.5 Plus21.9%estimated ± 0.7 pp, low confidence
32Claude Opus 4.521.6%estimated ± 0.7 pp, low confidence
33Grok 420.5%estimated ± 0.7 pp, low confidence
34o319.6%estimated ± 0.7 pp, low confidence
35GLM-517.4%estimated ± 0.7 pp, low confidence
36Gemini 2.5 Pro15.1%estimated ± 0.7 pp, low confidence
37Claude Sonnet 4.514.5%estimated ± 0.7 pp, low confidence
38o110.3%estimated ± 0.7 pp, low confidence
39Qwen3 235B 2507 (Reasoning)9.5%estimated ± 0.7 pp, low confidence
40Qwen3.5 Flash7.3%estimated ± 0.7 pp, low confidence
41Claude Haiku 4.57.0%estimated ± 0.7 pp, low confidence
42GPT-4.16.6%estimated ± 0.7 pp, low confidence
43Gemini 2.5 Flash5.9%estimated ± 0.7 pp, low confidence
44GPT-4.1 mini5.6%estimated ± 0.7 pp, low confidence
45GLM-4.64.9%estimated ± 0.7 pp, low confidence
46Grok 3 [Beta]4.9%estimated ± 0.7 pp, low confidence
47GLM-4.73.6%estimated ± 0.7 pp, low confidence
48Claude 3.5 Sonnet3.2%estimated ± 0.7 pp, low confidence
49DeepSeek V32.8%estimated ± 0.7 pp, low confidence
50GPT-4.1 nano2.2%estimated ± 0.7 pp, low confidence
51Llama 4 Maverick1.8%estimated ± 0.7 pp, low confidence
52GPT-4o1.5%estimated ± 0.7 pp, low confidence
53Llama 4 Scout1.1%estimated ± 0.7 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General