benchgap
Calibration

GPQA → C-Eval

C-Eval is estimated from GPQA with a Michaelis–Menten curve fitted on 5 models measured on both: y = 2.0000·x / (1.03238 + x), R² = 0.70, cross-validated error 0.9 pp. It is used for 38 estimates.

Estimated modelGPQAC-EvalSource
Claude 3.5 Sonnet59.4%73.0%estimated ± 0.9 pp, low confidence
Claude Sonnet 4.583.4%89.4%estimated ± 0.9 pp, low confidence
DeepSeek V359.1%72.8%estimated ± 0.9 pp, low confidence
Gemma 4 12B78.8%86.6%estimated ± 0.9 pp, low confidence
Gemma 4 E2B43.4%59.2%estimated ± 0.9 pp, low confidence
Gemma 4 E4B58.6%72.4%estimated ± 0.9 pp, low confidence
GPT-4.166.3%78.2%estimated ± 0.9 pp, low confidence
GPT-4.1 mini64.2%76.7%estimated ± 0.9 pp, low confidence
GPT-4.1 nano50.3%65.5%estimated ± 0.9 pp, low confidence
GPT-5.292.4%94.5%estimated ± 0.9 pp, low confidence
GPT-5.6 Luna92.3%94.4%estimated ± 0.9 pp, low confidence
GPT-5.6 Sol94.6%95.6%estimated ± 0.9 pp, low confidence
GPT-5.6 Terra92.9%94.7%estimated ± 0.9 pp, low confidence
GPT-6 Astra96.0%96.4%estimated ± 0.9 pp, low confidence
Granite 4.2 30B66.4%78.3%estimated ± 0.9 pp, low confidence
Granite 4.2 3B54.8%69.4%estimated ± 0.9 pp, low confidence
Granite 4.2 8B64.1%76.6%estimated ± 0.9 pp, low confidence
Interfaze Beta89.9%93.1%estimated ± 0.9 pp, medium confidence
Kimi K2.5 (Reasoning)87.6%91.8%estimated ± 0.9 pp, medium confidence
LFM2.5-230M25.4%39.5%estimated ± 0.9 pp, low confidence
LFM2.5-VL-450M25.7%39.8%estimated ± 0.9 pp, low confidence
Ling 2.6 Flash59.0%72.7%estimated ± 0.9 pp, low confidence
Ling 3.0 Flash FP884.0%89.7%estimated ± 0.9 pp, low confidence
MAI-Thinking-184.2%89.8%estimated ± 0.9 pp, low confidence
Mellum2-12B-A2.5B-Instruct40.9%56.8%estimated ± 0.9 pp, low confidence
Mellum2-12B-A2.5B-Thinking57.6%71.6%estimated ± 0.9 pp, low confidence
MiMo-V2-Flash83.7%89.5%estimated ± 0.9 pp, low confidence
Nemotron 3.5 Lightning 30B A3B NVFP475.6%84.5%estimated ± 0.9 pp, low confidence
Nemotron 3 Nano Omni 30B A3B72.2%82.3%estimated ± 0.9 pp, low confidence
o175.7%84.6%estimated ± 0.9 pp, low confidence
o1-pro79.0%86.7%estimated ± 0.9 pp, low confidence
o3-mini77.2%85.6%estimated ± 0.9 pp, low confidence
Sakana Fugu95.5%96.1%estimated ± 0.9 pp, low confidence
Sakana Fugu-Ultra95.5%96.1%estimated ± 0.9 pp, low confidence
Soofi S 30B-A3B43.4%59.2%estimated ± 0.9 pp, low confidence
Ternary Bonsai 2 27B85.8%90.8%estimated ± 0.9 pp, low confidence
ZAYA1-74B-Preview57.3%71.4%estimated ± 0.9 pp, low confidence
ZAYA1-8B71.0%81.5%estimated ± 0.9 pp, low confidence