benchgap
Calibration

GPQA → MMLU-Pro

MMLU-Pro is estimated from GPQA with a Hill curve fitted on 37 models measured on both: y = 0.0000 + (0.9486 − 0.0000)·x^2.83 / (0.39734^2.83 + x^2.83), R² = 0.97, cross-validated error 3.0 pp. It is used for 24 estimates.

Estimated modelGPQAMMLU-ProSource
Claude 3.5 Sonnet59.4%71.8%estimated ± 3.0 pp, high confidence
Claude Sonnet 4.583.4%84.5%estimated ± 3.0 pp, high confidence
DeepSeek V4.1 Flash90.9%86.5%estimated ± 3.0 pp, high confidence
Gemini 2.5 Pro83.0%84.4%estimated ± 3.0 pp, high confidence
Gemini 3.5 Flash92.2%86.8%estimated ± 3.0 pp, high confidence
GPT-4.166.3%76.8%estimated ± 3.0 pp, high confidence
GPT-4.1 mini64.2%75.5%estimated ± 3.0 pp, high confidence
GPT-4.1 nano50.3%62.7%estimated ± 3.0 pp, high confidence
GPT-5.292.4%86.9%estimated ± 3.0 pp, high confidence
GPT-5.6 Luna92.3%86.9%estimated ± 3.0 pp, high confidence
GPT-5.6 Sol94.6%87.4%estimated ± 3.0 pp, medium confidence
GPT-5.6 Terra92.9%87.0%estimated ± 3.0 pp, medium confidence
GPT-6 Astra96.0%87.6%estimated ± 3.0 pp, medium confidence
Grok 4.390.1%86.4%estimated ± 3.0 pp, high confidence
Hy3 Preview87.2%85.6%estimated ± 3.0 pp, high confidence
Interfaze Beta89.9%86.3%estimated ± 3.0 pp, high confidence
Kimi K2.690.5%86.5%estimated ± 3.0 pp, high confidence
Ling 2.6 Flash59.0%71.5%estimated ± 3.0 pp, high confidence
Ling 3.0 Flash85.0%85.0%estimated ± 3.0 pp, high confidence
Ling 3.0 Flash FP884.0%84.7%estimated ± 3.0 pp, high confidence
o175.7%81.7%estimated ± 3.0 pp, high confidence
o1-pro79.0%83.0%estimated ± 3.0 pp, high confidence
o3-mini77.2%82.3%estimated ± 3.0 pp, high confidence
Qwen3.8-Omni-Flash91.0%86.6%estimated ± 3.0 pp, high confidence