Calibration
FrontierMath v2 (Tier 4) → MMAnswerBench
MMAnswerBench is estimated from FrontierMath v2 (Tier 4) with a linear curve fitted on 6 models measured on both: y = 0.2259·x + 0.8192, R² = 0.62, cross-validated error 1.3 pp. It is used for 39 estimates.
| Estimated model | FrontierMath v2 (Tier 4) | MMAnswerBench | Source |
|---|---|---|---|
| Claude 3.5 Sonnet | 0.0% | 81.9% | estimated ± 1.3 pp, low confidence |
| Claude Haiku 4.5 | 2.1% | 82.4% | estimated ± 1.3 pp, low confidence |
| Claude Opus 4.6 | 22.9% | 87.1% | estimated ± 1.3 pp, low confidence |
| Claude Opus 4.7 | 22.9% | 87.1% | estimated ± 1.3 pp, low confidence |
| Claude Opus 4.8 | 31.3% | 89.0% | estimated ± 1.3 pp, low confidence |
| Claude Sonnet 4.5 | 4.2% | 82.9% | estimated ± 1.3 pp, medium confidence |
| Claude Sonnet 4.6 | 8.3% | 83.8% | estimated ± 1.3 pp, medium confidence |
| DeepSeek V3.2 | 2.1% | 82.4% | estimated ± 1.3 pp, medium confidence |
| Gemini 2.5 Flash | 4.2% | 82.9% | estimated ± 1.3 pp, medium confidence |
| Gemini 2.5 Pro | 4.2% | 82.9% | estimated ± 1.3 pp, medium confidence |
| Gemini 3.1 Pro | 16.7% | 85.7% | estimated ± 1.3 pp, low confidence |
| Gemini 3.5 Flash | 14.6% | 85.2% | estimated ± 1.3 pp, low confidence |
| Gemini 3 Flash | 4.2% | 82.9% | estimated ± 1.3 pp, medium confidence |
| Gemini 3 Pro | 18.8% | 86.2% | estimated ± 1.3 pp, low confidence |
| GLM-4.6 | 2.1% | 82.4% | estimated ± 1.3 pp, medium confidence |
| GLM-4.7 | 0.0% | 81.9% | estimated ± 1.3 pp, low confidence |
| GPT-4.1 | 0.0% | 81.9% | estimated ± 1.3 pp, low confidence |
| GPT-5.1 | 12.5% | 84.7% | estimated ± 1.3 pp, medium confidence |
| GPT-5.2 | 18.8% | 86.2% | estimated ± 1.3 pp, low confidence |
| GPT-5.4 | 27.1% | 88.0% | estimated ± 1.3 pp, low confidence |
| GPT-5.4 mini | 2.1% | 82.4% | estimated ± 1.3 pp, low confidence |
| GPT-5.4 nano | 6.3% | 83.3% | estimated ± 1.3 pp, medium confidence |
| GPT-5.4 Pro | 37.5% | 90.4% | estimated ± 1.3 pp, low confidence |
| GPT-5.5 | 35.4% | 89.9% | estimated ± 1.3 pp, low confidence |
| GPT-5.5 Pro | 39.6% | 90.9% | estimated ± 1.3 pp, low confidence |
| GPT-5.6 Luna | 58.5% | 95.1% | estimated ± 1.3 pp, low confidence |
| GPT-5.6 Sol | 83.0% | 100.0% | estimated ± 1.3 pp, low confidence |
| GPT-5.6 Terra | 68.3% | 97.4% | estimated ± 1.3 pp, low confidence |
| GPT-6 Astra | 97.6% | 100.0% | estimated ± 1.3 pp, low confidence |
| Grok 3 [Beta] | 0.0% | 81.9% | estimated ± 1.3 pp, low confidence |
| Grok 4 | 2.1% | 82.4% | estimated ± 1.3 pp, low confidence |
| Kimi K2 | 0.0% | 81.9% | estimated ± 1.3 pp, low confidence |
| Muse Spark | 14.6% | 85.2% | estimated ± 1.3 pp, low confidence |
| o3 | 2.1% | 82.4% | estimated ± 1.3 pp, low confidence |
| o4-mini (high) | 6.3% | 83.3% | estimated ± 1.3 pp, medium confidence |
| Qwen3 235B 2507 (Reasoning) | 0.0% | 81.9% | estimated ± 1.3 pp, low confidence |
| Qwen3.5 Flash | 0.0% | 81.9% | estimated ± 1.3 pp, low confidence |
| Qwen3.5 Plus | 2.1% | 82.4% | estimated ± 1.3 pp, low confidence |
| Qwen 3.6 Max (preview) | 4.2% | 82.9% | estimated ± 1.3 pp, medium confidence |