benchgap
Calibration

FrontierMath v2 (Tier 4) → MMAnswerBench

MMAnswerBench is estimated from FrontierMath v2 (Tier 4) with a linear curve fitted on 6 models measured on both: y = 0.2259·x + 0.8192, R² = 0.62, cross-validated error 1.3 pp. It is used for 39 estimates.

Estimated modelFrontierMath v2 (Tier 4)MMAnswerBenchSource
Claude 3.5 Sonnet0.0%81.9%estimated ± 1.3 pp, low confidence
Claude Haiku 4.52.1%82.4%estimated ± 1.3 pp, low confidence
Claude Opus 4.622.9%87.1%estimated ± 1.3 pp, low confidence
Claude Opus 4.722.9%87.1%estimated ± 1.3 pp, low confidence
Claude Opus 4.831.3%89.0%estimated ± 1.3 pp, low confidence
Claude Sonnet 4.54.2%82.9%estimated ± 1.3 pp, medium confidence
Claude Sonnet 4.68.3%83.8%estimated ± 1.3 pp, medium confidence
DeepSeek V3.22.1%82.4%estimated ± 1.3 pp, medium confidence
Gemini 2.5 Flash4.2%82.9%estimated ± 1.3 pp, medium confidence
Gemini 2.5 Pro4.2%82.9%estimated ± 1.3 pp, medium confidence
Gemini 3.1 Pro16.7%85.7%estimated ± 1.3 pp, low confidence
Gemini 3.5 Flash14.6%85.2%estimated ± 1.3 pp, low confidence
Gemini 3 Flash4.2%82.9%estimated ± 1.3 pp, medium confidence
Gemini 3 Pro18.8%86.2%estimated ± 1.3 pp, low confidence
GLM-4.62.1%82.4%estimated ± 1.3 pp, medium confidence
GLM-4.70.0%81.9%estimated ± 1.3 pp, low confidence
GPT-4.10.0%81.9%estimated ± 1.3 pp, low confidence
GPT-5.112.5%84.7%estimated ± 1.3 pp, medium confidence
GPT-5.218.8%86.2%estimated ± 1.3 pp, low confidence
GPT-5.427.1%88.0%estimated ± 1.3 pp, low confidence
GPT-5.4 mini2.1%82.4%estimated ± 1.3 pp, low confidence
GPT-5.4 nano6.3%83.3%estimated ± 1.3 pp, medium confidence
GPT-5.4 Pro37.5%90.4%estimated ± 1.3 pp, low confidence
GPT-5.535.4%89.9%estimated ± 1.3 pp, low confidence
GPT-5.5 Pro39.6%90.9%estimated ± 1.3 pp, low confidence
GPT-5.6 Luna58.5%95.1%estimated ± 1.3 pp, low confidence
GPT-5.6 Sol83.0%100.0%estimated ± 1.3 pp, low confidence
GPT-5.6 Terra68.3%97.4%estimated ± 1.3 pp, low confidence
GPT-6 Astra97.6%100.0%estimated ± 1.3 pp, low confidence
Grok 3 [Beta]0.0%81.9%estimated ± 1.3 pp, low confidence
Grok 42.1%82.4%estimated ± 1.3 pp, low confidence
Kimi K20.0%81.9%estimated ± 1.3 pp, low confidence
Muse Spark14.6%85.2%estimated ± 1.3 pp, low confidence
o32.1%82.4%estimated ± 1.3 pp, low confidence
o4-mini (high)6.3%83.3%estimated ± 1.3 pp, medium confidence
Qwen3 235B 2507 (Reasoning)0.0%81.9%estimated ± 1.3 pp, low confidence
Qwen3.5 Flash0.0%81.9%estimated ± 1.3 pp, low confidence
Qwen3.5 Plus2.1%82.4%estimated ± 1.3 pp, low confidence
Qwen 3.6 Max (preview)4.2%82.9%estimated ± 1.3 pp, medium confidence