benchgap
Calibration

GPQA Diamond → MMMLU

MMMLU is estimated from GPQA Diamond with a Hill curve fitted on 5 models measured on both: y = 0.0000 + (0.9548 − 0.0000)·x^6.00 / (0.56774^6.00 + x^6.00), R² = 0.93, cross-validated error 1.1 pp. It is used for 62 estimates.

Estimated modelGPQA DiamondMMMLUSource
A.X K285.6%88.0%estimated ± 1.1 pp, medium confidence
Claude Opus 4.689.2%89.5%estimated ± 1.1 pp, medium confidence
Claude Opus 4.7 (Adaptive)94.2%91.1%estimated ± 1.1 pp, low confidence
Claude Opus 4.893.6%90.9%estimated ± 1.1 pp, low confidence
DeepSeek V4.1 Flash90.9%90.1%estimated ± 1.1 pp, medium confidence
DeepSeek V4 Flash 073188.1%89.1%estimated ± 1.1 pp, medium confidence
DeepSeek V4 Pro 081390.1%89.9%estimated ± 1.1 pp, medium confidence
Gemini 3.1 Pro94.3%91.1%estimated ± 1.1 pp, low confidence
Gemini 3.5 Flash92.7%90.7%estimated ± 1.1 pp, low confidence
GLM-586.0%88.2%estimated ± 1.1 pp, medium confidence
GLM-5.186.2%88.3%estimated ± 1.1 pp, medium confidence
GLM-5.291.2%90.2%estimated ± 1.1 pp, medium confidence
GPT-5.492.8%90.7%estimated ± 1.1 pp, low confidence
GPT-5.593.6%90.9%estimated ± 1.1 pp, low confidence
GPT-5.6 Luna92.3%90.6%estimated ± 1.1 pp, medium confidence
GPT-5.6 Sol94.6%91.2%estimated ± 1.1 pp, low confidence
GPT-5.6 Terra92.9%90.8%estimated ± 1.1 pp, low confidence
GPT-6 Astra96.0%91.6%estimated ± 1.1 pp, low confidence
Grok 4.2088.5%89.3%estimated ± 1.1 pp, medium confidence
Hy3 Preview87.2%88.7%estimated ± 1.1 pp, medium confidence
Hy4 preview92.3%90.6%estimated ± 1.1 pp, medium confidence
Inkling87.9%89.0%estimated ± 1.1 pp, medium confidence
Inkling-Small89.5%89.6%estimated ± 1.1 pp, medium confidence
Kimi K2.690.5%90.0%estimated ± 1.1 pp, medium confidence
Kimi K2.587.6%88.9%estimated ± 1.1 pp, medium confidence
Kimi K393.5%90.9%estimated ± 1.1 pp, low confidence
LFM2.5-230M25.4%0.8%estimated ± 1.1 pp, low confidence
Ling 3.0 Flash85.0%87.7%estimated ± 1.1 pp, medium confidence
Ling 3.0 Flash FP884.0%87.2%estimated ± 1.1 pp, medium confidence
LLaDA2.2-mini44.4%17.8%estimated ± 1.1 pp, low confidence
LongCat-Flash-Lite-Sparse69.5%73.6%estimated ± 1.1 pp, low confidence
MAI-Thinking-184.2%87.3%estimated ± 1.1 pp, medium confidence
Mellum2-12B-A2.5B-Instruct40.9%11.7%estimated ± 1.1 pp, low confidence
Mellum2-12B-A2.5B-Thinking57.6%49.8%estimated ± 1.1 pp, low confidence
Mercury 2.579.0%83.9%estimated ± 1.1 pp, medium confidence
MiniCPM5-1B26.3%0.9%estimated ± 1.1 pp, low confidence
MiniCPM5-2B70.2%74.6%estimated ± 1.1 pp, low confidence
MiniMax M2.787.0%88.6%estimated ± 1.1 pp, medium confidence
Muse Spark89.5%89.6%estimated ± 1.1 pp, medium confidence
Nemotron 3.5 Lightning 30B A3B NVFP475.6%80.9%estimated ± 1.1 pp, low confidence
Nemotron 3 Nano Omni 30B A3B72.2%77.2%estimated ± 1.1 pp, low confidence
Nemotron 3 Ultra87.0%88.6%estimated ± 1.1 pp, medium confidence
Ornith-1.5-35B-A3B89.2%89.5%estimated ± 1.1 pp, medium confidence
Ornith-1.5-397B92.8%90.7%estimated ± 1.1 pp, low confidence
Ornith-1.5-9B86.4%88.4%estimated ± 1.1 pp, medium confidence
Pareto 26.10 Preview92.4%90.6%estimated ± 1.1 pp, medium confidence
Qwen3.8-27B89.2%89.5%estimated ± 1.1 pp, medium confidence
Qwen3.8-Flash-Next91.7%90.4%estimated ± 1.1 pp, medium confidence
Qwen3.8 Max92.6%90.7%estimated ± 1.1 pp, low confidence
Qwen3.8-Omni-Flash91.0%90.2%estimated ± 1.1 pp, medium confidence
Beam90.5%90.0%estimated ± 1.1 pp, medium confidence
Sakana Fugu95.5%91.4%estimated ± 1.1 pp, low confidence
Sakana Fugu-Ultra95.5%91.4%estimated ± 1.1 pp, low confidence
Solar Open 286.3%88.3%estimated ± 1.1 pp, medium confidence
Solar Pro 489.0%89.5%estimated ± 1.1 pp, medium confidence
Soofi S 30B-A3B43.4%15.9%estimated ± 1.1 pp, low confidence
Step 5 Preview93.5%90.9%estimated ± 1.1 pp, low confidence
Ternary Bonsai 2 27B85.8%88.1%estimated ± 1.1 pp, medium confidence
Trinity-Large-Preview63.3%62.8%estimated ± 1.1 pp, low confidence
Trinity-Large-Thinking76.3%81.6%estimated ± 1.1 pp, low confidence
ZAYA1-74B-Preview57.3%49.1%estimated ± 1.1 pp, low confidence
ZAYA1-8B71.0%75.7%estimated ± 1.1 pp, low confidence