benchgap
Calibration

GPQA Diamond → MMLU-Pro (Arcee)

MMLU-Pro (Arcee) is estimated from GPQA Diamond with a Michaelis–Menten curve fitted on 6 models measured on both: y = 1.3364·x / (0.48454 + x), R² = 0.72, cross-validated error 3.4 pp. It is used for 61 estimates.

Estimated modelGPQA DiamondMMLU-Pro (Arcee)Source
A.X K285.6%85.3%estimated ± 3.4 pp, medium confidence
Claude Opus 4.7 (Adaptive)94.2%88.2%estimated ± 3.4 pp, low confidence
Claude Opus 4.893.6%88.1%estimated ± 3.4 pp, low confidence
DeepSeek V4.1 Flash90.9%87.2%estimated ± 3.4 pp, low confidence
DeepSeek V4 Flash 073188.1%86.2%estimated ± 3.4 pp, medium confidence
DeepSeek V4 Pro 081390.1%86.9%estimated ± 3.4 pp, low confidence
Gemini 3.1 Pro94.3%88.3%estimated ± 3.4 pp, low confidence
Gemini 3.5 Flash92.7%87.8%estimated ± 3.4 pp, low confidence
Gemma 4 12B78.8%82.8%estimated ± 3.4 pp, medium confidence
GLM-5.186.2%85.6%estimated ± 3.4 pp, medium confidence
GLM-5.291.2%87.3%estimated ± 3.4 pp, low confidence
GPT-5.492.8%87.8%estimated ± 3.4 pp, low confidence
GPT-5.593.6%88.1%estimated ± 3.4 pp, low confidence
GPT-5.6 Luna92.3%87.6%estimated ± 3.4 pp, low confidence
GPT-5.6 Sol94.6%88.4%estimated ± 3.4 pp, low confidence
GPT-5.6 Terra92.9%87.8%estimated ± 3.4 pp, low confidence
GPT-6 Astra96.0%88.8%estimated ± 3.4 pp, low confidence
Grok 4.2088.5%86.4%estimated ± 3.4 pp, medium confidence
Hy3 Preview87.2%85.9%estimated ± 3.4 pp, medium confidence
Hy4 preview92.3%87.6%estimated ± 3.4 pp, low confidence
Inkling87.9%86.2%estimated ± 3.4 pp, medium confidence
Inkling-Small89.5%86.7%estimated ± 3.4 pp, low confidence
Interfaze Beta89.9%86.8%estimated ± 3.4 pp, low confidence
K-EXAONE 2.082.2%84.1%estimated ± 3.4 pp, medium confidence
Kimi K2.690.5%87.0%estimated ± 3.4 pp, low confidence
Kimi K393.5%88.0%estimated ± 3.4 pp, low confidence
LFM2.5-230M25.4%46.0%estimated ± 3.4 pp, low confidence
Ling 3.0 Flash85.0%85.1%estimated ± 3.4 pp, medium confidence
Ling 3.0 Flash FP884.0%84.8%estimated ± 3.4 pp, medium confidence
LLaDA2.2-mini44.4%63.9%estimated ± 3.4 pp, low confidence
LongCat-Flash-Lite-Sparse69.5%78.7%estimated ± 3.4 pp, medium confidence
MAI-Thinking-184.2%84.8%estimated ± 3.4 pp, medium confidence
Mellum2-12B-A2.5B-Instruct40.9%61.2%estimated ± 3.4 pp, low confidence
Mellum2-12B-A2.5B-Thinking57.6%72.6%estimated ± 3.4 pp, low confidence
Mercury 2.579.0%82.8%estimated ± 3.4 pp, medium confidence
MiniCPM5-1B26.3%47.0%estimated ± 3.4 pp, low confidence
MiniCPM5-2B70.2%79.1%estimated ± 3.4 pp, medium confidence
Muse Spark89.5%86.7%estimated ± 3.4 pp, low confidence
Nemotron 3.5 Lightning 30B A3B NVFP475.6%81.4%estimated ± 3.4 pp, medium confidence
Nemotron 3 Nano Omni 30B A3B72.2%80.0%estimated ± 3.4 pp, medium confidence
Nemotron 3 Ultra87.0%85.8%estimated ± 3.4 pp, medium confidence
Ornith-1.5-35B-A3B89.2%86.6%estimated ± 3.4 pp, medium confidence
Ornith-1.5-397B92.8%87.8%estimated ± 3.4 pp, low confidence
Ornith-1.5-9B86.4%85.6%estimated ± 3.4 pp, medium confidence
Pareto 26.10 Preview92.4%87.7%estimated ± 3.4 pp, low confidence
Qwen3.7 Max92.4%87.7%estimated ± 3.4 pp, low confidence
Qwen3.7 Plus90.3%87.0%estimated ± 3.4 pp, low confidence
Qwen3.8-27B89.2%86.6%estimated ± 3.4 pp, medium confidence
Qwen3.8-Flash-Next91.7%87.4%estimated ± 3.4 pp, low confidence
Qwen3.8 Max92.6%87.7%estimated ± 3.4 pp, low confidence
Qwen3.8-Omni-Flash91.0%87.2%estimated ± 3.4 pp, low confidence
Beam90.5%87.0%estimated ± 3.4 pp, low confidence
Sakana Fugu95.5%88.7%estimated ± 3.4 pp, low confidence
Sakana Fugu-Ultra95.5%88.7%estimated ± 3.4 pp, low confidence
Solar Open 286.3%85.6%estimated ± 3.4 pp, medium confidence
Solar Pro 489.0%86.5%estimated ± 3.4 pp, medium confidence
Soofi S 30B-A3B43.4%63.1%estimated ± 3.4 pp, low confidence
Step 5 Preview93.5%88.0%estimated ± 3.4 pp, low confidence
Ternary Bonsai 2 27B85.8%85.4%estimated ± 3.4 pp, medium confidence
ZAYA1-74B-Preview57.3%72.4%estimated ± 3.4 pp, low confidence
ZAYA1-8B71.0%79.4%estimated ± 3.4 pp, medium confidence