benchgap
Calibration

FrontierMath v2 (Tier 4) → HMMT Nov 2025

HMMT Nov 2025 is estimated from FrontierMath v2 (Tier 4) with a offset logistic curve fitted on 5 models measured on both: y = 1.0000 + (0.9322 − 1.0000) / (1 + exp(−200.00·(x − 0.0214))), R² = 0.58, cross-validated error 2.6 pp. It is used for 39 estimates.

Estimated modelFrontierMath v2 (Tier 4)HMMT Nov 2025Source
Claude 3.5 Sonnet0.0%99.9%estimated ± 2.6 pp, low confidence
Claude Haiku 4.52.1%96.8%estimated ± 2.6 pp, low confidence
Claude Opus 4.622.9%93.2%estimated ± 2.6 pp, low confidence
Claude Opus 4.722.9%93.2%estimated ± 2.6 pp, low confidence
Claude Opus 4.831.3%93.2%estimated ± 2.6 pp, low confidence
Claude Sonnet 4.54.2%93.3%estimated ± 2.6 pp, medium confidence
Claude Sonnet 4.68.3%93.2%estimated ± 2.6 pp, medium confidence
DeepSeek V3.22.1%96.8%estimated ± 2.6 pp, medium confidence
Gemini 2.5 Flash4.2%93.3%estimated ± 2.6 pp, medium confidence
Gemini 2.5 Pro4.2%93.3%estimated ± 2.6 pp, medium confidence
Gemini 3.1 Pro16.7%93.2%estimated ± 2.6 pp, low confidence
Gemini 3.5 Flash14.6%93.2%estimated ± 2.6 pp, low confidence
Gemini 3 Flash4.2%93.3%estimated ± 2.6 pp, medium confidence
Gemini 3 Pro18.8%93.2%estimated ± 2.6 pp, low confidence
GLM-4.62.1%96.7%estimated ± 2.6 pp, medium confidence
GLM-4.70.0%99.9%estimated ± 2.6 pp, low confidence
GPT-4.10.0%99.9%estimated ± 2.6 pp, low confidence
GPT-5.112.5%93.2%estimated ± 2.6 pp, medium confidence
GPT-5.218.8%93.2%estimated ± 2.6 pp, low confidence
GPT-5.427.1%93.2%estimated ± 2.6 pp, low confidence
GPT-5.4 mini2.1%96.8%estimated ± 2.6 pp, low confidence
GPT-5.4 nano6.3%93.2%estimated ± 2.6 pp, medium confidence
GPT-5.4 Pro37.5%93.2%estimated ± 2.6 pp, low confidence
GPT-5.535.4%93.2%estimated ± 2.6 pp, low confidence
GPT-5.5 Pro39.6%93.2%estimated ± 2.6 pp, low confidence
GPT-5.6 Luna58.5%93.2%estimated ± 2.6 pp, low confidence
GPT-5.6 Sol83.0%93.2%estimated ± 2.6 pp, low confidence
GPT-5.6 Terra68.3%93.2%estimated ± 2.6 pp, low confidence
GPT-6 Astra97.6%93.2%estimated ± 2.6 pp, low confidence
Grok 3 [Beta]0.0%99.9%estimated ± 2.6 pp, low confidence
Grok 42.1%96.8%estimated ± 2.6 pp, low confidence
Kimi K20.0%99.9%estimated ± 2.6 pp, low confidence
Muse Spark14.6%93.2%estimated ± 2.6 pp, low confidence
o32.1%96.8%estimated ± 2.6 pp, low confidence
o4-mini (high)6.3%93.2%estimated ± 2.6 pp, medium confidence
Qwen3 235B 2507 (Reasoning)0.0%99.9%estimated ± 2.6 pp, low confidence
Qwen3.5 Flash0.0%99.9%estimated ± 2.6 pp, low confidence
Qwen3.5 Plus2.1%96.8%estimated ± 2.6 pp, low confidence
Qwen 3.6 Max (preview)4.2%93.3%estimated ± 2.6 pp, medium confidence