benchgap
Calibration

FrontierMath v2 (Tier 4) → AIME26

AIME26 is estimated from FrontierMath v2 (Tier 4) with a offset logistic curve fitted on 6 models measured on both: y = 0.9546 + (1.2000 − 0.9546) / (1 + exp(−200.00·(x − 0.1619))), R² = 0.64, cross-validated error 0.6 pp. It is used for 39 estimates.

Estimated modelFrontierMath v2 (Tier 4)AIME26Source
Claude 3.5 Sonnet0.0%95.5%estimated ± 0.6 pp, low confidence
Claude Haiku 4.52.1%95.5%estimated ± 0.6 pp, low confidence
Claude Opus 4.622.9%100.0%estimated ± 0.6 pp, low confidence
Claude Opus 4.722.9%100.0%estimated ± 0.6 pp, low confidence
Claude Opus 4.831.3%100.0%estimated ± 0.6 pp, low confidence
Claude Sonnet 4.54.2%95.5%estimated ± 0.6 pp, medium confidence
Claude Sonnet 4.68.3%95.5%estimated ± 0.6 pp, medium confidence
DeepSeek V3.22.1%95.5%estimated ± 0.6 pp, medium confidence
Gemini 2.5 Flash4.2%95.5%estimated ± 0.6 pp, medium confidence
Gemini 2.5 Pro4.2%95.5%estimated ± 0.6 pp, medium confidence
Gemini 3.1 Pro16.7%100.0%estimated ± 0.6 pp, low confidence
Gemini 3.5 Flash14.6%96.4%estimated ± 0.6 pp, low confidence
Gemini 3 Flash4.2%95.5%estimated ± 0.6 pp, medium confidence
Gemini 3 Pro18.8%100.0%estimated ± 0.6 pp, low confidence
GLM-4.62.1%95.5%estimated ± 0.6 pp, medium confidence
GLM-4.70.0%95.5%estimated ± 0.6 pp, low confidence
GPT-4.10.0%95.5%estimated ± 0.6 pp, low confidence
GPT-5.112.5%95.5%estimated ± 0.6 pp, medium confidence
GPT-5.218.8%100.0%estimated ± 0.6 pp, low confidence
GPT-5.427.1%100.0%estimated ± 0.6 pp, low confidence
GPT-5.4 mini2.1%95.5%estimated ± 0.6 pp, low confidence
GPT-5.4 nano6.3%95.5%estimated ± 0.6 pp, medium confidence
GPT-5.4 Pro37.5%100.0%estimated ± 0.6 pp, low confidence
GPT-5.535.4%100.0%estimated ± 0.6 pp, low confidence
GPT-5.5 Pro39.6%100.0%estimated ± 0.6 pp, low confidence
GPT-5.6 Luna58.5%100.0%estimated ± 0.6 pp, low confidence
GPT-5.6 Sol83.0%100.0%estimated ± 0.6 pp, low confidence
GPT-5.6 Terra68.3%100.0%estimated ± 0.6 pp, low confidence
GPT-6 Astra97.6%100.0%estimated ± 0.6 pp, low confidence
Grok 3 [Beta]0.0%95.5%estimated ± 0.6 pp, low confidence
Grok 42.1%95.5%estimated ± 0.6 pp, low confidence
Kimi K20.0%95.5%estimated ± 0.6 pp, low confidence
Muse Spark14.6%96.4%estimated ± 0.6 pp, low confidence
o32.1%95.5%estimated ± 0.6 pp, low confidence
o4-mini (high)6.3%95.5%estimated ± 0.6 pp, medium confidence
Qwen3 235B 2507 (Reasoning)0.0%95.5%estimated ± 0.6 pp, low confidence
Qwen3.5 Flash0.0%95.5%estimated ± 0.6 pp, low confidence
Qwen3.5 Plus2.1%95.5%estimated ± 0.6 pp, low confidence
Qwen 3.6 Max (preview)4.2%95.5%estimated ± 0.6 pp, medium confidence