benchgap
Calibration

AA-HLE → Vals MMLU-Pro

Vals MMLU-Pro is estimated from AA-HLE with a offset logistic curve fitted on 47 models measured on both: y = 0.8305 + (0.9324 − 0.8305) / (1 + exp(−13.26·(x − 0.4329))), R² = 0.50, cross-validated error 2.5 pp. It is used for 9 estimates.

Estimated modelAA-HLEVals MMLU-ProSource
Claude 3 Opus2.8%83.1%estimated ± 2.5 pp, medium confidence
Claude 4.1 Opus Thinking12.5%83.2%estimated ± 2.5 pp, high confidence
DeepSeek R1 Distill Qwen 32B4.6%83.1%estimated ± 2.5 pp, medium confidence
Gemini 1.0 Pro4.2%83.1%estimated ± 2.5 pp, medium confidence
Gemini 1.5 Pro4.6%83.1%estimated ± 2.5 pp, medium confidence
GPT-4 Turbo3.1%83.1%estimated ± 2.5 pp, medium confidence
GPT-4o mini4.2%83.1%estimated ± 2.5 pp, medium confidence
Phi-4 Multimodal Instruct5.0%83.1%estimated ± 2.5 pp, medium confidence
Qwen2.5 Coder 32B Instruct3.5%83.1%estimated ± 2.5 pp, medium confidence