benchgap
Calibration

Artificial Analysis Intelligence Index → MMLU

MMLU is estimated from Artificial Analysis Intelligence Index with a Michaelis–Menten curve fitted on 6 models measured on both: y = 1.0705·x / (0.02522 + x), R² = 0.89, cross-validated error 1.9 pp. It is used for 15 estimates.

Estimated modelArtificial Analysis Intelligence IndexMMLUSource
Claude 3 Opus8.7%83.0%estimated ± 1.9 pp, medium confidence
Claude 4.1 Opus18.6%94.2%estimated ± 1.9 pp, low confidence
Claude 4.1 Opus Thinking22.9%96.4%estimated ± 1.9 pp, low confidence
Claude Haiku 5.543.4%100.0%estimated ± 1.9 pp, low confidence
DeepSeek R1 Distill Qwen 32B8.4%82.3%estimated ± 1.9 pp, medium confidence
Gemini 1.0 Pro5.3%72.7%estimated ± 1.9 pp, low confidence
Gemini 1.5 Pro7.9%81.2%estimated ± 1.9 pp, medium confidence
GLM-5.3-Flash41.8%100.0%estimated ± 1.9 pp, low confidence
GPT-4 Turbo7.0%78.8%estimated ± 1.9 pp, low confidence
GPT-4o mini6.7%77.6%estimated ± 1.9 pp, low confidence
o1-preview11.4%87.6%estimated ± 1.9 pp, medium confidence
o1-pro12.4%89.0%estimated ± 1.9 pp, medium confidence
o3-pro21.9%96.0%estimated ± 1.9 pp, low confidence
Phi-4 Multimodal Instruct5.8%74.6%estimated ± 1.9 pp, low confidence
Qwen2.5 Coder 32B Instruct6.7%77.9%estimated ± 1.9 pp, low confidence