benchgap
Calibration

HLE w/o tools → MMLU-Pro

MMLU-Pro is estimated from HLE w/o tools with a Michaelis–Menten curve fitted on 7 models measured on both: y = 0.8685·x / (0.00587 + x), R² = 0.69, cross-validated error 2.5 pp. It is used for 35 estimates.

Estimated modelHLE w/o toolsMMLU-ProSource
Claude Fable 5.160.9%86.0%estimated ± 2.5 pp, low confidence
Claude Haiku 5.545.9%85.8%estimated ± 2.5 pp, low confidence
Claude Mythos 559.0%86.0%estimated ± 2.5 pp, low confidence
Claude Opus 4.7 (Adaptive)46.9%85.8%estimated ± 2.5 pp, low confidence
Claude Opus 4.849.8%85.8%estimated ± 2.5 pp, low confidence
Claude Opus 556.3%86.0%estimated ± 2.5 pp, low confidence
Claude Opus 5.564.4%86.1%estimated ± 2.5 pp, low confidence
Claude Sonnet 543.2%85.7%estimated ± 2.5 pp, low confidence
Claude Sonnet 5.556.9%86.0%estimated ± 2.5 pp, low confidence
Gemini 3.1 Pro45.4%85.7%estimated ± 2.5 pp, low confidence
GLM-5.240.5%85.6%estimated ± 2.5 pp, low confidence
GPT-5.439.8%85.6%estimated ± 2.5 pp, medium confidence
GPT-5.4 mini28.2%85.1%estimated ± 2.5 pp, medium confidence
GPT-5.4 nano24.3%84.8%estimated ± 2.5 pp, medium confidence
GPT-5.4 Pro42.7%85.7%estimated ± 2.5 pp, low confidence
GPT-5.541.4%85.6%estimated ± 2.5 pp, low confidence
GPT-5.5 Pro43.1%85.7%estimated ± 2.5 pp, low confidence
Grok 4.2031.6%85.3%estimated ± 2.5 pp, medium confidence
Hy4 preview43.4%85.7%estimated ± 2.5 pp, low confidence
Inkling30.0%85.2%estimated ± 2.5 pp, medium confidence
Inkling-Small31.6%85.3%estimated ± 2.5 pp, medium confidence
Kimi K343.5%85.7%estimated ± 2.5 pp, low confidence
MiMo-V2.5-Pro34.0%85.4%estimated ± 2.5 pp, medium confidence
Muse Spark42.8%85.7%estimated ± 2.5 pp, low confidence
Muse Spark 1.152.2%85.9%estimated ± 2.5 pp, low confidence
Ornith-1.5-35B-A3B25.6%84.9%estimated ± 2.5 pp, medium confidence
Ornith-1.5-397B44.6%85.7%estimated ± 2.5 pp, low confidence
Ornith-1.5-9B20.2%84.4%estimated ± 2.5 pp, medium confidence
Pareto 26.949.0%85.8%estimated ± 2.5 pp, low confidence
Qwen3.8-27B30.8%85.2%estimated ± 2.5 pp, medium confidence
Qwen3.8-Flash-Next35.9%85.5%estimated ± 2.5 pp, medium confidence
Qwen3.8 Max43.6%85.7%estimated ± 2.5 pp, low confidence
Beam36.2%85.5%estimated ± 2.5 pp, medium confidence
Sakana Fugu47.2%85.8%estimated ± 2.5 pp, low confidence
Sakana Fugu-Ultra50.0%85.8%estimated ± 2.5 pp, low confidence