Calibration
HLE → C-Eval
C-Eval is estimated from HLE with a Hill curve fitted on 5 models measured on both: y = 0.0084 + (0.9330 − 0.0084)·x^6.00 / (0.12370^6.00 + x^6.00), R² = 0.86, cross-validated error 0.8 pp. It is used for 47 estimates.
| Estimated model | HLE | C-Eval | Source |
|---|---|---|---|
| A.X K2 | 27.8% | 92.6% | estimated ± 0.8 pp, medium confidence |
| Agents-A1 | 47.6% | 93.3% | estimated ± 0.8 pp, low confidence |
| Apodex 1.1 | 56.1% | 93.3% | estimated ± 0.8 pp, low confidence |
| Claude Fable 5.1 | 65.0% | 93.3% | estimated ± 0.8 pp, low confidence |
| Claude Mythos 5 | 64.5% | 93.3% | estimated ± 0.8 pp, low confidence |
| Claude Opus 4.7 (Adaptive) | 54.7% | 93.3% | estimated ± 0.8 pp, low confidence |
| Claude Opus 4.8 | 57.9% | 93.3% | estimated ± 0.8 pp, low confidence |
| Claude Opus 5 | 64.7% | 93.3% | estimated ± 0.8 pp, low confidence |
| Claude Sonnet 5 | 57.4% | 93.3% | estimated ± 0.8 pp, low confidence |
| DeepSeek V4.1 Flash | 36.8% | 93.2% | estimated ± 0.8 pp, low confidence |
| DeepSeek V4 Flash 0731 | 34.8% | 93.1% | estimated ± 0.8 pp, low confidence |
| DeepSeek V4 Pro 0813 | 42.7% | 93.2% | estimated ± 0.8 pp, low confidence |
| dots3-note Preview | 52.6% | 93.3% | estimated ± 0.8 pp, low confidence |
| Gemini 2.5 Pro | 18.8% | 86.4% | estimated ± 0.8 pp, low confidence |
| Gemini 3.5 Flash | 40.2% | 93.2% | estimated ± 0.8 pp, low confidence |
| Gemma 4 26B A4B | 17.2% | 82.1% | estimated ± 0.8 pp, low confidence |
| Gemma 4 31B | 26.5% | 92.3% | estimated ± 0.8 pp, medium confidence |
| GLM-4.7 | 24.8% | 91.9% | estimated ± 0.8 pp, medium confidence |
| GLM-5.1 | 52.3% | 93.3% | estimated ± 0.8 pp, low confidence |
| GLM-5.2 | 54.7% | 93.3% | estimated ± 0.8 pp, low confidence |
| GPT-5.4 | 52.1% | 93.3% | estimated ± 0.8 pp, low confidence |
| GPT-5.4 mini | 41.5% | 93.2% | estimated ± 0.8 pp, low confidence |
| GPT-5.4 nano | 37.7% | 93.2% | estimated ± 0.8 pp, low confidence |
| GPT-5.4 Pro | 58.7% | 93.3% | estimated ± 0.8 pp, low confidence |
| GPT-5.5 | 52.2% | 93.3% | estimated ± 0.8 pp, low confidence |
| GPT-5.5 Pro | 57.2% | 93.3% | estimated ± 0.8 pp, low confidence |
| Grok 4.3 | 35.0% | 93.1% | estimated ± 0.8 pp, low confidence |
| Hy3 Preview | 25.5% | 92.1% | estimated ± 0.8 pp, medium confidence |
| Hy4 preview | 55.4% | 93.3% | estimated ± 0.8 pp, low confidence |
| Inkling | 46.0% | 93.3% | estimated ± 0.8 pp, low confidence |
| Inkling-Small | 47.8% | 93.3% | estimated ± 0.8 pp, low confidence |
| K-EXAONE 2.0 | 18.3% | 85.2% | estimated ± 0.8 pp, low confidence |
| Kimi K2.6 | 34.7% | 93.1% | estimated ± 0.8 pp, low confidence |
| Kimi K3 | 56.0% | 93.3% | estimated ± 0.8 pp, low confidence |
| Ling 3.0 Flash | 22.7% | 90.9% | estimated ± 0.8 pp, medium confidence |
| MiMo-V2.5-Pro | 48.0% | 93.3% | estimated ± 0.8 pp, low confidence |
| Muse Spark | 50.4% | 93.3% | estimated ± 0.8 pp, low confidence |
| Muse Spark 1.1 | 62.1% | 93.3% | estimated ± 0.8 pp, low confidence |
| Nemotron 3 Ultra | 26.7% | 92.4% | estimated ± 0.8 pp, medium confidence |
| Ornith-1.5-35B-A3B | 25.6% | 92.1% | estimated ± 0.8 pp, medium confidence |
| Ornith-1.5-397B | 44.6% | 93.3% | estimated ± 0.8 pp, low confidence |
| Ornith-1.5-9B | 20.2% | 88.7% | estimated ± 0.8 pp, low confidence |
| Qwen3.8-27B | 30.8% | 92.9% | estimated ± 0.8 pp, medium confidence |
| Qwen3.8-Flash-Next | 35.9% | 93.1% | estimated ± 0.8 pp, low confidence |
| Qwen3.8 Max | 43.6% | 93.2% | estimated ± 0.8 pp, low confidence |
| Qwen3.8-Omni-Flash | 36.5% | 93.2% | estimated ± 0.8 pp, low confidence |
| Step 5 Preview | 46.5% | 93.3% | estimated ± 0.8 pp, low confidence |