Calibration
GPQA → C-Eval
C-Eval is estimated from GPQA with a Michaelis–Menten curve fitted on 5 models measured on both: y = 2.0000·x / (1.03238 + x), R² = 0.70, cross-validated error 0.9 pp. It is used for 38 estimates.
| Estimated model | GPQA | C-Eval | Source |
|---|---|---|---|
| Claude 3.5 Sonnet | 59.4% | 73.0% | estimated ± 0.9 pp, low confidence |
| Claude Sonnet 4.5 | 83.4% | 89.4% | estimated ± 0.9 pp, low confidence |
| DeepSeek V3 | 59.1% | 72.8% | estimated ± 0.9 pp, low confidence |
| Gemma 4 12B | 78.8% | 86.6% | estimated ± 0.9 pp, low confidence |
| Gemma 4 E2B | 43.4% | 59.2% | estimated ± 0.9 pp, low confidence |
| Gemma 4 E4B | 58.6% | 72.4% | estimated ± 0.9 pp, low confidence |
| GPT-4.1 | 66.3% | 78.2% | estimated ± 0.9 pp, low confidence |
| GPT-4.1 mini | 64.2% | 76.7% | estimated ± 0.9 pp, low confidence |
| GPT-4.1 nano | 50.3% | 65.5% | estimated ± 0.9 pp, low confidence |
| GPT-5.2 | 92.4% | 94.5% | estimated ± 0.9 pp, low confidence |
| GPT-5.6 Luna | 92.3% | 94.4% | estimated ± 0.9 pp, low confidence |
| GPT-5.6 Sol | 94.6% | 95.6% | estimated ± 0.9 pp, low confidence |
| GPT-5.6 Terra | 92.9% | 94.7% | estimated ± 0.9 pp, low confidence |
| GPT-6 Astra | 96.0% | 96.4% | estimated ± 0.9 pp, low confidence |
| Granite 4.2 30B | 66.4% | 78.3% | estimated ± 0.9 pp, low confidence |
| Granite 4.2 3B | 54.8% | 69.4% | estimated ± 0.9 pp, low confidence |
| Granite 4.2 8B | 64.1% | 76.6% | estimated ± 0.9 pp, low confidence |
| Interfaze Beta | 89.9% | 93.1% | estimated ± 0.9 pp, medium confidence |
| Kimi K2.5 (Reasoning) | 87.6% | 91.8% | estimated ± 0.9 pp, medium confidence |
| LFM2.5-230M | 25.4% | 39.5% | estimated ± 0.9 pp, low confidence |
| LFM2.5-VL-450M | 25.7% | 39.8% | estimated ± 0.9 pp, low confidence |
| Ling 2.6 Flash | 59.0% | 72.7% | estimated ± 0.9 pp, low confidence |
| Ling 3.0 Flash FP8 | 84.0% | 89.7% | estimated ± 0.9 pp, low confidence |
| MAI-Thinking-1 | 84.2% | 89.8% | estimated ± 0.9 pp, low confidence |
| Mellum2-12B-A2.5B-Instruct | 40.9% | 56.8% | estimated ± 0.9 pp, low confidence |
| Mellum2-12B-A2.5B-Thinking | 57.6% | 71.6% | estimated ± 0.9 pp, low confidence |
| MiMo-V2-Flash | 83.7% | 89.5% | estimated ± 0.9 pp, low confidence |
| Nemotron 3.5 Lightning 30B A3B NVFP4 | 75.6% | 84.5% | estimated ± 0.9 pp, low confidence |
| Nemotron 3 Nano Omni 30B A3B | 72.2% | 82.3% | estimated ± 0.9 pp, low confidence |
| o1 | 75.7% | 84.6% | estimated ± 0.9 pp, low confidence |
| o1-pro | 79.0% | 86.7% | estimated ± 0.9 pp, low confidence |
| o3-mini | 77.2% | 85.6% | estimated ± 0.9 pp, low confidence |
| Sakana Fugu | 95.5% | 96.1% | estimated ± 0.9 pp, low confidence |
| Sakana Fugu-Ultra | 95.5% | 96.1% | estimated ± 0.9 pp, low confidence |
| Soofi S 30B-A3B | 43.4% | 59.2% | estimated ± 0.9 pp, low confidence |
| Ternary Bonsai 2 27B | 85.8% | 90.8% | estimated ± 0.9 pp, low confidence |
| ZAYA1-74B-Preview | 57.3% | 71.4% | estimated ± 0.9 pp, low confidence |
| ZAYA1-8B | 71.0% | 81.5% | estimated ± 0.9 pp, low confidence |