benchgap
Calibration

SuperGPQA → C-Eval

C-Eval is estimated from SuperGPQA with a Michaelis–Menten curve fitted on 5 models measured on both: y = 1.3120·x / (0.29241 + x), R² = 0.89, cross-validated error 0.7 pp. It is used for 15 estimates.

Estimated modelSuperGPQAC-EvalSource
Claude Opus 4.695.0%100.0%estimated ± 0.7 pp, low confidence
Claude Sonnet 4.695.0%100.0%estimated ± 0.7 pp, low confidence
GLM-566.8%91.3%estimated ± 0.7 pp, medium confidence
Kimi K2.569.2%92.2%estimated ± 0.7 pp, medium confidence
MiniCPM5-1B23.1%58.0%estimated ± 0.7 pp, low confidence
MiniCPM5-2B40.8%76.4%estimated ± 0.7 pp, low confidence
Qwen3 235B 250762.6%89.4%estimated ± 0.7 pp, low confidence
Qwen3.5-122B-A10B67.1%91.4%estimated ± 0.7 pp, medium confidence
Qwen3.5-27B65.6%90.7%estimated ± 0.7 pp, medium confidence
Qwen3.5-35B-A3B63.4%89.8%estimated ± 0.7 pp, low confidence
Qwen 3.6 Max (preview)73.9%94.0%estimated ± 0.7 pp, low confidence
Qwen3.7 Max73.6%93.9%estimated ± 0.7 pp, low confidence
Qwen3.7 Plus71.4%93.1%estimated ± 0.7 pp, medium confidence
Seed 2.1 Pro70.8%92.8%estimated ± 0.7 pp, medium confidence
Seed 2.1 Turbo67.4%91.5%estimated ± 0.7 pp, medium confidence