Calibration
FrontierMath v2 (Tier 4) → AIME26
AIME26 is estimated from FrontierMath v2 (Tier 4) with a offset logistic curve fitted on 6 models measured on both: y = 0.9546 + (1.2000 − 0.9546) / (1 + exp(−200.00·(x − 0.1619))), R² = 0.64, cross-validated error 0.6 pp. It is used for 39 estimates.
| Estimated model | FrontierMath v2 (Tier 4) | AIME26 | Source |
|---|---|---|---|
| Claude 3.5 Sonnet | 0.0% | 95.5% | estimated ± 0.6 pp, low confidence |
| Claude Haiku 4.5 | 2.1% | 95.5% | estimated ± 0.6 pp, low confidence |
| Claude Opus 4.6 | 22.9% | 100.0% | estimated ± 0.6 pp, low confidence |
| Claude Opus 4.7 | 22.9% | 100.0% | estimated ± 0.6 pp, low confidence |
| Claude Opus 4.8 | 31.3% | 100.0% | estimated ± 0.6 pp, low confidence |
| Claude Sonnet 4.5 | 4.2% | 95.5% | estimated ± 0.6 pp, medium confidence |
| Claude Sonnet 4.6 | 8.3% | 95.5% | estimated ± 0.6 pp, medium confidence |
| DeepSeek V3.2 | 2.1% | 95.5% | estimated ± 0.6 pp, medium confidence |
| Gemini 2.5 Flash | 4.2% | 95.5% | estimated ± 0.6 pp, medium confidence |
| Gemini 2.5 Pro | 4.2% | 95.5% | estimated ± 0.6 pp, medium confidence |
| Gemini 3.1 Pro | 16.7% | 100.0% | estimated ± 0.6 pp, low confidence |
| Gemini 3.5 Flash | 14.6% | 96.4% | estimated ± 0.6 pp, low confidence |
| Gemini 3 Flash | 4.2% | 95.5% | estimated ± 0.6 pp, medium confidence |
| Gemini 3 Pro | 18.8% | 100.0% | estimated ± 0.6 pp, low confidence |
| GLM-4.6 | 2.1% | 95.5% | estimated ± 0.6 pp, medium confidence |
| GLM-4.7 | 0.0% | 95.5% | estimated ± 0.6 pp, low confidence |
| GPT-4.1 | 0.0% | 95.5% | estimated ± 0.6 pp, low confidence |
| GPT-5.1 | 12.5% | 95.5% | estimated ± 0.6 pp, medium confidence |
| GPT-5.2 | 18.8% | 100.0% | estimated ± 0.6 pp, low confidence |
| GPT-5.4 | 27.1% | 100.0% | estimated ± 0.6 pp, low confidence |
| GPT-5.4 mini | 2.1% | 95.5% | estimated ± 0.6 pp, low confidence |
| GPT-5.4 nano | 6.3% | 95.5% | estimated ± 0.6 pp, medium confidence |
| GPT-5.4 Pro | 37.5% | 100.0% | estimated ± 0.6 pp, low confidence |
| GPT-5.5 | 35.4% | 100.0% | estimated ± 0.6 pp, low confidence |
| GPT-5.5 Pro | 39.6% | 100.0% | estimated ± 0.6 pp, low confidence |
| GPT-5.6 Luna | 58.5% | 100.0% | estimated ± 0.6 pp, low confidence |
| GPT-5.6 Sol | 83.0% | 100.0% | estimated ± 0.6 pp, low confidence |
| GPT-5.6 Terra | 68.3% | 100.0% | estimated ± 0.6 pp, low confidence |
| GPT-6 Astra | 97.6% | 100.0% | estimated ± 0.6 pp, low confidence |
| Grok 3 [Beta] | 0.0% | 95.5% | estimated ± 0.6 pp, low confidence |
| Grok 4 | 2.1% | 95.5% | estimated ± 0.6 pp, low confidence |
| Kimi K2 | 0.0% | 95.5% | estimated ± 0.6 pp, low confidence |
| Muse Spark | 14.6% | 96.4% | estimated ± 0.6 pp, low confidence |
| o3 | 2.1% | 95.5% | estimated ± 0.6 pp, low confidence |
| o4-mini (high) | 6.3% | 95.5% | estimated ± 0.6 pp, medium confidence |
| Qwen3 235B 2507 (Reasoning) | 0.0% | 95.5% | estimated ± 0.6 pp, low confidence |
| Qwen3.5 Flash | 0.0% | 95.5% | estimated ± 0.6 pp, low confidence |
| Qwen3.5 Plus | 2.1% | 95.5% | estimated ± 0.6 pp, low confidence |
| Qwen 3.6 Max (preview) | 4.2% | 95.5% | estimated ± 0.6 pp, medium confidence |