benchgap
Calibration

FrontierMath v2 (Tiers 1-3) → FrontierMath (legacy)

FrontierMath (legacy) is estimated from FrontierMath v2 (Tiers 1-3) with a linear curve fitted on 6 models measured on both: y = 0.9865·x + 0.0114, R² = 1.00, cross-validated error 0.7 pp. It is used for 45 estimates.

Estimated modelFrontierMath v2 (Tiers 1-3)FrontierMath (legacy)Source
Claude 3.5 Sonnet2.1%3.2%estimated ± 0.7 pp, low confidence
Claude Haiku 4.55.9%7.0%estimated ± 0.7 pp, low confidence
Claude Opus 4.520.7%21.6%estimated ± 0.7 pp, low confidence
Claude Opus 4.640.7%41.3%estimated ± 0.7 pp, low confidence
Claude Opus 4.743.8%44.3%estimated ± 0.7 pp, low confidence
Claude Opus 4.847.2%47.7%estimated ± 0.7 pp, low confidence
Claude Sonnet 4.513.5%14.5%estimated ± 0.7 pp, low confidence
Claude Sonnet 4.632.4%33.1%estimated ± 0.7 pp, low confidence
DeepSeek V31.7%2.8%estimated ± 0.7 pp, low confidence
DeepSeek V3.222.1%22.9%estimated ± 0.7 pp, low confidence
Gemini 2.5 Flash4.8%5.9%estimated ± 0.7 pp, low confidence
Gemini 2.5 Pro14.1%15.1%estimated ± 0.7 pp, low confidence
Gemini 3.1 Pro36.9%37.5%estimated ± 0.7 pp, low confidence
Gemini 3.5 Flash39.0%39.6%estimated ± 0.7 pp, low confidence
Gemini 3 Flash35.6%36.3%estimated ± 0.7 pp, low confidence
Gemini 3 Pro37.6%38.2%estimated ± 0.7 pp, low confidence
GLM-4.63.8%4.9%estimated ± 0.7 pp, low confidence
GLM-4.72.4%3.6%estimated ± 0.7 pp, low confidence
GLM-516.4%17.4%estimated ± 0.7 pp, low confidence
GLM-5.133.4%34.1%estimated ± 0.7 pp, low confidence
GPT-4.15.5%6.6%estimated ± 0.7 pp, low confidence
GPT-4.1 mini4.5%5.6%estimated ± 0.7 pp, low confidence
GPT-4.1 nano1.0%2.2%estimated ± 0.7 pp, low confidence
GPT-4o0.3%1.5%estimated ± 0.7 pp, low confidence
GPT-5.131.0%31.8%estimated ± 0.7 pp, low confidence
GPT-5.240.7%41.3%estimated ± 0.7 pp, low confidence
GPT-5.447.6%48.1%estimated ± 0.7 pp, low confidence
GPT-5.4 mini28.3%29.0%estimated ± 0.7 pp, low confidence
GPT-5.4 nano25.9%26.7%estimated ± 0.7 pp, low confidence
Grok 3 [Beta]3.8%4.9%estimated ± 0.7 pp, low confidence
Grok 419.7%20.5%estimated ± 0.7 pp, low confidence
Kimi K2.639.0%39.6%estimated ± 0.7 pp, low confidence
Kimi K221.4%22.3%estimated ± 0.7 pp, low confidence
Kimi K2.527.9%28.7%estimated ± 0.7 pp, low confidence
Llama 4 Maverick0.7%1.8%estimated ± 0.7 pp, low confidence
Llama 4 Scout0.0%1.1%estimated ± 0.7 pp, low confidence
Muse Spark39.0%39.6%estimated ± 0.7 pp, low confidence
o19.3%10.3%estimated ± 0.7 pp, low confidence
o318.7%19.6%estimated ± 0.7 pp, low confidence
o4-mini (high)24.8%25.6%estimated ± 0.7 pp, low confidence
Qwen3 235B 2507 (Reasoning)8.5%9.5%estimated ± 0.7 pp, low confidence
Qwen3.5 Flash6.2%7.3%estimated ± 0.7 pp, low confidence
Qwen3.5 Plus21.0%21.9%estimated ± 0.7 pp, low confidence
Qwen 3.6 Max (preview)23.1%23.9%estimated ± 0.7 pp, low confidence
Qwen3.6 Plus26.2%27.0%estimated ± 0.7 pp, low confidence