Calibration
FrontierMath v2 (Tier 4) → HMMT Nov 2025
HMMT Nov 2025 is estimated from FrontierMath v2 (Tier 4) with a offset logistic curve fitted on 5 models measured on both: y = 1.0000 + (0.9322 − 1.0000) / (1 + exp(−200.00·(x − 0.0214))), R² = 0.58, cross-validated error 2.6 pp. It is used for 39 estimates.
| Estimated model | FrontierMath v2 (Tier 4) | HMMT Nov 2025 | Source |
|---|---|---|---|
| Claude 3.5 Sonnet | 0.0% | 99.9% | estimated ± 2.6 pp, low confidence |
| Claude Haiku 4.5 | 2.1% | 96.8% | estimated ± 2.6 pp, low confidence |
| Claude Opus 4.6 | 22.9% | 93.2% | estimated ± 2.6 pp, low confidence |
| Claude Opus 4.7 | 22.9% | 93.2% | estimated ± 2.6 pp, low confidence |
| Claude Opus 4.8 | 31.3% | 93.2% | estimated ± 2.6 pp, low confidence |
| Claude Sonnet 4.5 | 4.2% | 93.3% | estimated ± 2.6 pp, medium confidence |
| Claude Sonnet 4.6 | 8.3% | 93.2% | estimated ± 2.6 pp, medium confidence |
| DeepSeek V3.2 | 2.1% | 96.8% | estimated ± 2.6 pp, medium confidence |
| Gemini 2.5 Flash | 4.2% | 93.3% | estimated ± 2.6 pp, medium confidence |
| Gemini 2.5 Pro | 4.2% | 93.3% | estimated ± 2.6 pp, medium confidence |
| Gemini 3.1 Pro | 16.7% | 93.2% | estimated ± 2.6 pp, low confidence |
| Gemini 3.5 Flash | 14.6% | 93.2% | estimated ± 2.6 pp, low confidence |
| Gemini 3 Flash | 4.2% | 93.3% | estimated ± 2.6 pp, medium confidence |
| Gemini 3 Pro | 18.8% | 93.2% | estimated ± 2.6 pp, low confidence |
| GLM-4.6 | 2.1% | 96.7% | estimated ± 2.6 pp, medium confidence |
| GLM-4.7 | 0.0% | 99.9% | estimated ± 2.6 pp, low confidence |
| GPT-4.1 | 0.0% | 99.9% | estimated ± 2.6 pp, low confidence |
| GPT-5.1 | 12.5% | 93.2% | estimated ± 2.6 pp, medium confidence |
| GPT-5.2 | 18.8% | 93.2% | estimated ± 2.6 pp, low confidence |
| GPT-5.4 | 27.1% | 93.2% | estimated ± 2.6 pp, low confidence |
| GPT-5.4 mini | 2.1% | 96.8% | estimated ± 2.6 pp, low confidence |
| GPT-5.4 nano | 6.3% | 93.2% | estimated ± 2.6 pp, medium confidence |
| GPT-5.4 Pro | 37.5% | 93.2% | estimated ± 2.6 pp, low confidence |
| GPT-5.5 | 35.4% | 93.2% | estimated ± 2.6 pp, low confidence |
| GPT-5.5 Pro | 39.6% | 93.2% | estimated ± 2.6 pp, low confidence |
| GPT-5.6 Luna | 58.5% | 93.2% | estimated ± 2.6 pp, low confidence |
| GPT-5.6 Sol | 83.0% | 93.2% | estimated ± 2.6 pp, low confidence |
| GPT-5.6 Terra | 68.3% | 93.2% | estimated ± 2.6 pp, low confidence |
| GPT-6 Astra | 97.6% | 93.2% | estimated ± 2.6 pp, low confidence |
| Grok 3 [Beta] | 0.0% | 99.9% | estimated ± 2.6 pp, low confidence |
| Grok 4 | 2.1% | 96.8% | estimated ± 2.6 pp, low confidence |
| Kimi K2 | 0.0% | 99.9% | estimated ± 2.6 pp, low confidence |
| Muse Spark | 14.6% | 93.2% | estimated ± 2.6 pp, low confidence |
| o3 | 2.1% | 96.8% | estimated ± 2.6 pp, low confidence |
| o4-mini (high) | 6.3% | 93.2% | estimated ± 2.6 pp, medium confidence |
| Qwen3 235B 2507 (Reasoning) | 0.0% | 99.9% | estimated ± 2.6 pp, low confidence |
| Qwen3.5 Flash | 0.0% | 99.9% | estimated ± 2.6 pp, low confidence |
| Qwen3.5 Plus | 2.1% | 96.8% | estimated ± 2.6 pp, low confidence |
| Qwen 3.6 Max (preview) | 4.2% | 93.3% | estimated ± 2.6 pp, medium confidence |