Calibration
HLE → MMLU-Redux
MMLU-Redux is estimated from HLE with a Michaelis–Menten + offset curve fitted on 7 models measured on both: y = 0.0000 + 0.9959·x / (0.01541 + x), R² = 0.95, cross-validated error 2.1 pp. It is used for 11 estimates.
| Estimated model | HLE | MMLU-Redux | Source |
|---|---|---|---|
| Agents-A1 | 47.6% | 96.5% | estimated ± 2.1 pp, low confidence |
| Claude Mythos 5 | 64.5% | 97.3% | estimated ± 2.1 pp, low confidence |
| dots3-note Preview | 52.6% | 96.8% | estimated ± 2.1 pp, low confidence |
| GPT-5.4 Pro | 58.7% | 97.0% | estimated ± 2.1 pp, low confidence |
| GPT-5.5 Pro | 57.2% | 97.0% | estimated ± 2.1 pp, low confidence |
| Hy4 preview | 55.4% | 96.9% | estimated ± 2.1 pp, low confidence |
| Ornith-1.5-35B-A3B | 25.6% | 93.9% | estimated ± 2.1 pp, medium confidence |
| Ornith-1.5-397B | 44.6% | 96.3% | estimated ± 2.1 pp, low confidence |
| Ornith-1.5-9B | 20.2% | 92.5% | estimated ± 2.1 pp, medium confidence |
| Qwen3.8 Max | 43.6% | 96.2% | estimated ± 2.1 pp, low confidence |
| Qwen3.8-Omni-Flash | 36.5% | 95.6% | estimated ± 2.1 pp, medium confidence |