Calibration
Artificial Analysis Intelligence Index → MMLU
MMLU is estimated from Artificial Analysis Intelligence Index with a Michaelis–Menten curve fitted on 6 models measured on both: y = 1.0705·x / (0.02522 + x), R² = 0.89, cross-validated error 1.9 pp. It is used for 15 estimates.
| Estimated model | Artificial Analysis Intelligence Index | MMLU | Source |
|---|---|---|---|
| Claude 3 Opus | 8.7% | 83.0% | estimated ± 1.9 pp, medium confidence |
| Claude 4.1 Opus | 18.6% | 94.2% | estimated ± 1.9 pp, low confidence |
| Claude 4.1 Opus Thinking | 22.9% | 96.4% | estimated ± 1.9 pp, low confidence |
| Claude Haiku 5.5 | 43.4% | 100.0% | estimated ± 1.9 pp, low confidence |
| DeepSeek R1 Distill Qwen 32B | 8.4% | 82.3% | estimated ± 1.9 pp, medium confidence |
| Gemini 1.0 Pro | 5.3% | 72.7% | estimated ± 1.9 pp, low confidence |
| Gemini 1.5 Pro | 7.9% | 81.2% | estimated ± 1.9 pp, medium confidence |
| GLM-5.3-Flash | 41.8% | 100.0% | estimated ± 1.9 pp, low confidence |
| GPT-4 Turbo | 7.0% | 78.8% | estimated ± 1.9 pp, low confidence |
| GPT-4o mini | 6.7% | 77.6% | estimated ± 1.9 pp, low confidence |
| o1-preview | 11.4% | 87.6% | estimated ± 1.9 pp, medium confidence |
| o1-pro | 12.4% | 89.0% | estimated ± 1.9 pp, medium confidence |
| o3-pro | 21.9% | 96.0% | estimated ± 1.9 pp, low confidence |
| Phi-4 Multimodal Instruct | 5.8% | 74.6% | estimated ± 1.9 pp, low confidence |
| Qwen2.5 Coder 32B Instruct | 6.7% | 77.9% | estimated ± 1.9 pp, low confidence |