benchgap
Calibration

GPQA → Vals MMLU-Pro

Vals MMLU-Pro is estimated from GPQA with a offset logistic curve fitted on 24 models measured on both: y = 0.0000 + (0.8959 − 0.0000) / (1 + exp(−22.32·(x − 0.7447))), R² = 0.86, cross-validated error 1.2 pp. It is used for 60 estimates.

Estimated modelGPQAVals MMLU-ProSource
Claude 3.5 Sonnet59.4%3.0%estimated ± 1.2 pp, medium confidence
Claude Mythos 594.1%88.5%estimated ± 1.2 pp, high confidence
Claude Opus 4.587.0%84.4%estimated ± 1.2 pp, high confidence
Claude Opus 4.691.3%87.5%estimated ± 1.2 pp, high confidence
Claude Opus 4.7 (Adaptive)94.2%88.5%estimated ± 1.2 pp, high confidence
Claude Sonnet 4.583.4%78.8%estimated ± 1.2 pp, high confidence
DeepSeek V359.1%2.8%estimated ± 1.2 pp, medium confidence
DeepSeek V4.1 Flash90.9%87.4%estimated ± 1.2 pp, high confidence
Gemini 2.5 Pro83.0%78.0%estimated ± 1.2 pp, high confidence
Gemma 4 12B78.8%64.9%estimated ± 1.2 pp, medium confidence
Gemma 4 31B84.3%80.6%estimated ± 1.2 pp, high confidence
Gemma 4 E2B43.4%0.1%estimated ± 1.2 pp, medium confidence
Gemma 4 E4B58.6%2.5%estimated ± 1.2 pp, medium confidence
GLM-586.0%83.2%estimated ± 1.2 pp, high confidence
GPT-4.166.3%12.4%estimated ± 1.2 pp, medium confidence
GPT-4.1 mini64.2%8.2%estimated ± 1.2 pp, medium confidence
GPT-4.1 nano50.3%0.4%estimated ± 1.2 pp, medium confidence
GPT-5.292.4%88.0%estimated ± 1.2 pp, high confidence
GPT-5.492.8%88.1%estimated ± 1.2 pp, high confidence
GPT-6 Astra96.0%88.9%estimated ± 1.2 pp, medium confidence
Granite 4.2 30B66.4%12.7%estimated ± 1.2 pp, medium confidence
Granite 4.2 3B54.8%1.1%estimated ± 1.2 pp, medium confidence
Granite 4.2 8B64.1%8.1%estimated ± 1.2 pp, medium confidence
Hy3 Preview87.2%84.6%estimated ± 1.2 pp, high confidence
Hy4 preview92.3%87.9%estimated ± 1.2 pp, high confidence
Interfaze Beta89.9%86.8%estimated ± 1.2 pp, high confidence
Kimi K2.587.6%85.0%estimated ± 1.2 pp, high confidence
Kimi K2.5 (Reasoning)87.6%85.0%estimated ± 1.2 pp, high confidence
LFM2.5-230M25.4%0.0%estimated ± 1.2 pp, medium confidence
LFM2.5-VL-450M25.7%0.0%estimated ± 1.2 pp, medium confidence
Ling 2.6 Flash59.0%2.7%estimated ± 1.2 pp, medium confidence
Ling 3.0 Flash FP884.0%80.0%estimated ± 1.2 pp, high confidence
MAI-Thinking-184.2%80.4%estimated ± 1.2 pp, high confidence
Mellum2-12B-A2.5B-Instruct40.9%0.0%estimated ± 1.2 pp, medium confidence
Mellum2-12B-A2.5B-Thinking57.6%2.0%estimated ± 1.2 pp, medium confidence
MiMo-V2-Flash83.7%79.5%estimated ± 1.2 pp, high confidence
Nemotron 3.5 Lightning 30B A3B NVFP475.6%50.2%estimated ± 1.2 pp, medium confidence
Nemotron 3 Nano Omni 30B A3B72.2%33.7%estimated ± 1.2 pp, medium confidence
o175.7%50.9%estimated ± 1.2 pp, medium confidence
o1-pro79.0%65.7%estimated ± 1.2 pp, medium confidence
o3-mini77.2%58.0%estimated ± 1.2 pp, medium confidence
Ornith-1.5-35B-A3B89.2%86.4%estimated ± 1.2 pp, high confidence
Ornith-1.5-397B92.8%88.1%estimated ± 1.2 pp, high confidence
Ornith-1.5-9B86.4%83.7%estimated ± 1.2 pp, high confidence
Qwen3 235B 250777.5%59.4%estimated ± 1.2 pp, medium confidence
Qwen3.5-122B-A10B86.6%84.0%estimated ± 1.2 pp, high confidence
Qwen3.5-27B85.5%82.5%estimated ± 1.2 pp, high confidence
Qwen3.5-35B-A3B84.2%80.4%estimated ± 1.2 pp, high confidence
Qwen3.5 397B88.4%85.8%estimated ± 1.2 pp, high confidence
Qwen3.6-27B87.8%85.2%estimated ± 1.2 pp, high confidence
Qwen3.6-35B-A3B86.0%83.2%estimated ± 1.2 pp, high confidence
Qwen3.7 Plus90.3%87.0%estimated ± 1.2 pp, high confidence
Qwen3.8-Flash-Next91.7%87.7%estimated ± 1.2 pp, high confidence
Qwen3.8-Omni-Flash91.0%87.4%estimated ± 1.2 pp, high confidence
Sakana Fugu95.5%88.8%estimated ± 1.2 pp, medium confidence
Sakana Fugu-Ultra95.5%88.8%estimated ± 1.2 pp, medium confidence
Soofi S 30B-A3B43.4%0.1%estimated ± 1.2 pp, medium confidence
Ternary Bonsai 2 27B85.8%82.9%estimated ± 1.2 pp, high confidence
ZAYA1-74B-Preview57.3%1.9%estimated ± 1.2 pp, medium confidence
ZAYA1-8B71.0%28.2%estimated ± 1.2 pp, medium confidence