Calibration
FrontierMath v2 (Tiers 1-3) → MMAnswerBench
MMAnswerBench is estimated from FrontierMath v2 (Tiers 1-3) with a Hill curve fitted on 6 models measured on both: y = 0.8282 + (1.2000 − 0.8282)·x^6.00 / (0.58642^6.00 + x^6.00), R² = 0.61, cross-validated error 1.5 pp. It is used for 7 estimates.
| Estimated model | FrontierMath v2 (Tiers 1-3) | MMAnswerBench | Source |
|---|---|---|---|
| DeepSeek V3 | 1.7% | 82.8% | estimated ± 1.5 pp, low confidence |
| GPT-4.1 mini | 4.5% | 82.8% | estimated ± 1.5 pp, low confidence |
| GPT-4.1 nano | 1.0% | 82.8% | estimated ± 1.5 pp, low confidence |
| GPT-4o | 0.3% | 82.8% | estimated ± 1.5 pp, low confidence |
| Llama 4 Maverick | 0.7% | 82.8% | estimated ± 1.5 pp, low confidence |
| Llama 4 Scout | 0.0% | 82.8% | estimated ± 1.5 pp, low confidence |
| o1 | 9.3% | 82.8% | estimated ± 1.5 pp, low confidence |