Calibration
FrontierMath v2 (Tiers 1-3) → FrontierMath (legacy)
FrontierMath (legacy) is estimated from FrontierMath v2 (Tiers 1-3) with a linear curve fitted on 6 models measured on both: y = 0.9865·x + 0.0114, R² = 1.00, cross-validated error 0.7 pp. It is used for 45 estimates.
| Estimated model | FrontierMath v2 (Tiers 1-3) | FrontierMath (legacy) | Source |
|---|---|---|---|
| Claude 3.5 Sonnet | 2.1% | 3.2% | estimated ± 0.7 pp, low confidence |
| Claude Haiku 4.5 | 5.9% | 7.0% | estimated ± 0.7 pp, low confidence |
| Claude Opus 4.5 | 20.7% | 21.6% | estimated ± 0.7 pp, low confidence |
| Claude Opus 4.6 | 40.7% | 41.3% | estimated ± 0.7 pp, low confidence |
| Claude Opus 4.7 | 43.8% | 44.3% | estimated ± 0.7 pp, low confidence |
| Claude Opus 4.8 | 47.2% | 47.7% | estimated ± 0.7 pp, low confidence |
| Claude Sonnet 4.5 | 13.5% | 14.5% | estimated ± 0.7 pp, low confidence |
| Claude Sonnet 4.6 | 32.4% | 33.1% | estimated ± 0.7 pp, low confidence |
| DeepSeek V3 | 1.7% | 2.8% | estimated ± 0.7 pp, low confidence |
| DeepSeek V3.2 | 22.1% | 22.9% | estimated ± 0.7 pp, low confidence |
| Gemini 2.5 Flash | 4.8% | 5.9% | estimated ± 0.7 pp, low confidence |
| Gemini 2.5 Pro | 14.1% | 15.1% | estimated ± 0.7 pp, low confidence |
| Gemini 3.1 Pro | 36.9% | 37.5% | estimated ± 0.7 pp, low confidence |
| Gemini 3.5 Flash | 39.0% | 39.6% | estimated ± 0.7 pp, low confidence |
| Gemini 3 Flash | 35.6% | 36.3% | estimated ± 0.7 pp, low confidence |
| Gemini 3 Pro | 37.6% | 38.2% | estimated ± 0.7 pp, low confidence |
| GLM-4.6 | 3.8% | 4.9% | estimated ± 0.7 pp, low confidence |
| GLM-4.7 | 2.4% | 3.6% | estimated ± 0.7 pp, low confidence |
| GLM-5 | 16.4% | 17.4% | estimated ± 0.7 pp, low confidence |
| GLM-5.1 | 33.4% | 34.1% | estimated ± 0.7 pp, low confidence |
| GPT-4.1 | 5.5% | 6.6% | estimated ± 0.7 pp, low confidence |
| GPT-4.1 mini | 4.5% | 5.6% | estimated ± 0.7 pp, low confidence |
| GPT-4.1 nano | 1.0% | 2.2% | estimated ± 0.7 pp, low confidence |
| GPT-4o | 0.3% | 1.5% | estimated ± 0.7 pp, low confidence |
| GPT-5.1 | 31.0% | 31.8% | estimated ± 0.7 pp, low confidence |
| GPT-5.2 | 40.7% | 41.3% | estimated ± 0.7 pp, low confidence |
| GPT-5.4 | 47.6% | 48.1% | estimated ± 0.7 pp, low confidence |
| GPT-5.4 mini | 28.3% | 29.0% | estimated ± 0.7 pp, low confidence |
| GPT-5.4 nano | 25.9% | 26.7% | estimated ± 0.7 pp, low confidence |
| Grok 3 [Beta] | 3.8% | 4.9% | estimated ± 0.7 pp, low confidence |
| Grok 4 | 19.7% | 20.5% | estimated ± 0.7 pp, low confidence |
| Kimi K2.6 | 39.0% | 39.6% | estimated ± 0.7 pp, low confidence |
| Kimi K2 | 21.4% | 22.3% | estimated ± 0.7 pp, low confidence |
| Kimi K2.5 | 27.9% | 28.7% | estimated ± 0.7 pp, low confidence |
| Llama 4 Maverick | 0.7% | 1.8% | estimated ± 0.7 pp, low confidence |
| Llama 4 Scout | 0.0% | 1.1% | estimated ± 0.7 pp, low confidence |
| Muse Spark | 39.0% | 39.6% | estimated ± 0.7 pp, low confidence |
| o1 | 9.3% | 10.3% | estimated ± 0.7 pp, low confidence |
| o3 | 18.7% | 19.6% | estimated ± 0.7 pp, low confidence |
| o4-mini (high) | 24.8% | 25.6% | estimated ± 0.7 pp, low confidence |
| Qwen3 235B 2507 (Reasoning) | 8.5% | 9.5% | estimated ± 0.7 pp, low confidence |
| Qwen3.5 Flash | 6.2% | 7.3% | estimated ± 0.7 pp, low confidence |
| Qwen3.5 Plus | 21.0% | 21.9% | estimated ± 0.7 pp, low confidence |
| Qwen 3.6 Max (preview) | 23.1% | 23.9% | estimated ± 0.7 pp, low confidence |
| Qwen3.6 Plus | 26.2% | 27.0% | estimated ± 0.7 pp, low confidence |