Calibration
GPQA → MMLU-Pro
MMLU-Pro is estimated from GPQA with a Hill curve fitted on 37 models measured on both: y = 0.0000 + (0.9486 − 0.0000)·x^2.83 / (0.39734^2.83 + x^2.83), R² = 0.97, cross-validated error 3.0 pp. It is used for 24 estimates.
| Estimated model | GPQA | MMLU-Pro | Source |
|---|---|---|---|
| Claude 3.5 Sonnet | 59.4% | 71.8% | estimated ± 3.0 pp, high confidence |
| Claude Sonnet 4.5 | 83.4% | 84.5% | estimated ± 3.0 pp, high confidence |
| DeepSeek V4.1 Flash | 90.9% | 86.5% | estimated ± 3.0 pp, high confidence |
| Gemini 2.5 Pro | 83.0% | 84.4% | estimated ± 3.0 pp, high confidence |
| Gemini 3.5 Flash | 92.2% | 86.8% | estimated ± 3.0 pp, high confidence |
| GPT-4.1 | 66.3% | 76.8% | estimated ± 3.0 pp, high confidence |
| GPT-4.1 mini | 64.2% | 75.5% | estimated ± 3.0 pp, high confidence |
| GPT-4.1 nano | 50.3% | 62.7% | estimated ± 3.0 pp, high confidence |
| GPT-5.2 | 92.4% | 86.9% | estimated ± 3.0 pp, high confidence |
| GPT-5.6 Luna | 92.3% | 86.9% | estimated ± 3.0 pp, high confidence |
| GPT-5.6 Sol | 94.6% | 87.4% | estimated ± 3.0 pp, medium confidence |
| GPT-5.6 Terra | 92.9% | 87.0% | estimated ± 3.0 pp, medium confidence |
| GPT-6 Astra | 96.0% | 87.6% | estimated ± 3.0 pp, medium confidence |
| Grok 4.3 | 90.1% | 86.4% | estimated ± 3.0 pp, high confidence |
| Hy3 Preview | 87.2% | 85.6% | estimated ± 3.0 pp, high confidence |
| Interfaze Beta | 89.9% | 86.3% | estimated ± 3.0 pp, high confidence |
| Kimi K2.6 | 90.5% | 86.5% | estimated ± 3.0 pp, high confidence |
| Ling 2.6 Flash | 59.0% | 71.5% | estimated ± 3.0 pp, high confidence |
| Ling 3.0 Flash | 85.0% | 85.0% | estimated ± 3.0 pp, high confidence |
| Ling 3.0 Flash FP8 | 84.0% | 84.7% | estimated ± 3.0 pp, high confidence |
| o1 | 75.7% | 81.7% | estimated ± 3.0 pp, high confidence |
| o1-pro | 79.0% | 83.0% | estimated ± 3.0 pp, high confidence |
| o3-mini | 77.2% | 82.3% | estimated ± 3.0 pp, high confidence |
| Qwen3.8-Omni-Flash | 91.0% | 86.6% | estimated ± 3.0 pp, high confidence |