benchgap
Calibration

GPQA → MMLU

MMLU is estimated from GPQA with a Hill curve fitted on 5 models measured on both: y = 0.0000 + (0.9076 − 0.0000)·x^6.00 / (0.35840^6.00 + x^6.00), R² = 0.80, cross-validated error 3.6 pp. It is used for 22 estimates.

Estimated modelGPQAMMLUSource
Claude 3.5 Sonnet59.4%86.6%estimated ± 3.6 pp, medium confidence
Claude Mythos 594.1%90.5%estimated ± 3.6 pp, low confidence
Claude Sonnet 4.583.4%90.2%estimated ± 3.6 pp, low confidence
Hy4 preview92.3%90.4%estimated ± 3.6 pp, low confidence
Interfaze Beta89.9%90.4%estimated ± 3.6 pp, low confidence
LFM2.5-230M25.4%10.2%estimated ± 3.6 pp, low confidence
LFM2.5-VL-450M25.7%10.8%estimated ± 3.6 pp, low confidence
MAI-Thinking-184.2%90.2%estimated ± 3.6 pp, low confidence
Mellum2-12B-A2.5B-Instruct40.9%62.5%estimated ± 3.6 pp, low confidence
Mellum2-12B-A2.5B-Thinking57.6%85.8%estimated ± 3.6 pp, medium confidence
Ornith-1.5-35B-A3B89.2%90.4%estimated ± 3.6 pp, low confidence
Ornith-1.5-397B92.8%90.5%estimated ± 3.6 pp, low confidence
Ornith-1.5-9B86.4%90.3%estimated ± 3.6 pp, low confidence
Qwen3 235B 250777.5%89.9%estimated ± 3.6 pp, low confidence
Qwen3.8 Max92.6%90.5%estimated ± 3.6 pp, low confidence
Qwen3.8-Omni-Flash91.0%90.4%estimated ± 3.6 pp, low confidence
Sakana Fugu95.5%90.5%estimated ± 3.6 pp, low confidence
Sakana Fugu-Ultra95.5%90.5%estimated ± 3.6 pp, low confidence
Soofi S 30B-A3B43.4%68.9%estimated ± 3.6 pp, low confidence
Ternary Bonsai 2 27B85.8%90.3%estimated ± 3.6 pp, low confidence
ZAYA1-74B-Preview57.3%85.6%estimated ± 3.6 pp, medium confidence
ZAYA1-8B71.0%89.3%estimated ± 3.6 pp, medium confidence