benchgap
Calibration

HLE → C-Eval

C-Eval is estimated from HLE with a Hill curve fitted on 5 models measured on both: y = 0.0084 + (0.9330 − 0.0084)·x^6.00 / (0.12370^6.00 + x^6.00), R² = 0.86, cross-validated error 0.8 pp. It is used for 47 estimates.

Estimated modelHLEC-EvalSource
A.X K227.8%92.6%estimated ± 0.8 pp, medium confidence
Agents-A147.6%93.3%estimated ± 0.8 pp, low confidence
Apodex 1.156.1%93.3%estimated ± 0.8 pp, low confidence
Claude Fable 5.165.0%93.3%estimated ± 0.8 pp, low confidence
Claude Mythos 564.5%93.3%estimated ± 0.8 pp, low confidence
Claude Opus 4.7 (Adaptive)54.7%93.3%estimated ± 0.8 pp, low confidence
Claude Opus 4.857.9%93.3%estimated ± 0.8 pp, low confidence
Claude Opus 564.7%93.3%estimated ± 0.8 pp, low confidence
Claude Sonnet 557.4%93.3%estimated ± 0.8 pp, low confidence
DeepSeek V4.1 Flash36.8%93.2%estimated ± 0.8 pp, low confidence
DeepSeek V4 Flash 073134.8%93.1%estimated ± 0.8 pp, low confidence
DeepSeek V4 Pro 081342.7%93.2%estimated ± 0.8 pp, low confidence
dots3-note Preview52.6%93.3%estimated ± 0.8 pp, low confidence
Gemini 2.5 Pro18.8%86.4%estimated ± 0.8 pp, low confidence
Gemini 3.5 Flash40.2%93.2%estimated ± 0.8 pp, low confidence
Gemma 4 26B A4B17.2%82.1%estimated ± 0.8 pp, low confidence
Gemma 4 31B26.5%92.3%estimated ± 0.8 pp, medium confidence
GLM-4.724.8%91.9%estimated ± 0.8 pp, medium confidence
GLM-5.152.3%93.3%estimated ± 0.8 pp, low confidence
GLM-5.254.7%93.3%estimated ± 0.8 pp, low confidence
GPT-5.452.1%93.3%estimated ± 0.8 pp, low confidence
GPT-5.4 mini41.5%93.2%estimated ± 0.8 pp, low confidence
GPT-5.4 nano37.7%93.2%estimated ± 0.8 pp, low confidence
GPT-5.4 Pro58.7%93.3%estimated ± 0.8 pp, low confidence
GPT-5.552.2%93.3%estimated ± 0.8 pp, low confidence
GPT-5.5 Pro57.2%93.3%estimated ± 0.8 pp, low confidence
Grok 4.335.0%93.1%estimated ± 0.8 pp, low confidence
Hy3 Preview25.5%92.1%estimated ± 0.8 pp, medium confidence
Hy4 preview55.4%93.3%estimated ± 0.8 pp, low confidence
Inkling46.0%93.3%estimated ± 0.8 pp, low confidence
Inkling-Small47.8%93.3%estimated ± 0.8 pp, low confidence
K-EXAONE 2.018.3%85.2%estimated ± 0.8 pp, low confidence
Kimi K2.634.7%93.1%estimated ± 0.8 pp, low confidence
Kimi K356.0%93.3%estimated ± 0.8 pp, low confidence
Ling 3.0 Flash22.7%90.9%estimated ± 0.8 pp, medium confidence
MiMo-V2.5-Pro48.0%93.3%estimated ± 0.8 pp, low confidence
Muse Spark50.4%93.3%estimated ± 0.8 pp, low confidence
Muse Spark 1.162.1%93.3%estimated ± 0.8 pp, low confidence
Nemotron 3 Ultra26.7%92.4%estimated ± 0.8 pp, medium confidence
Ornith-1.5-35B-A3B25.6%92.1%estimated ± 0.8 pp, medium confidence
Ornith-1.5-397B44.6%93.3%estimated ± 0.8 pp, low confidence
Ornith-1.5-9B20.2%88.7%estimated ± 0.8 pp, low confidence
Qwen3.8-27B30.8%92.9%estimated ± 0.8 pp, medium confidence
Qwen3.8-Flash-Next35.9%93.1%estimated ± 0.8 pp, low confidence
Qwen3.8 Max43.6%93.2%estimated ± 0.8 pp, low confidence
Qwen3.8-Omni-Flash36.5%93.2%estimated ± 0.8 pp, low confidence
Step 5 Preview46.5%93.3%estimated ± 0.8 pp, low confidence