benchgap
Calibration

Artificial Analysis Intelligence Index → SuperGPQA

SuperGPQA is estimated from Artificial Analysis Intelligence Index with a Michaelis–Menten curve fitted on 16 models measured on both: y = 1.3448·x / (0.20445 + x), R² = 0.44, cross-validated error 10.4 pp. It is used for 14 estimates.

Estimated modelArtificial Analysis Intelligence IndexSuperGPQASource
Claude 3 Opus8.7%40.2%estimated ± 10.4 pp, low confidence
Claude 4.1 Opus18.6%64.0%estimated ± 10.4 pp, low confidence
Claude 4.1 Opus Thinking22.9%71.0%estimated ± 10.4 pp, low confidence
Claude Haiku 5.543.4%91.4%estimated ± 10.4 pp, low confidence
DeepSeek R1 Distill Qwen 32B8.4%39.1%estimated ± 10.4 pp, low confidence
Gemini 1.0 Pro5.3%27.9%estimated ± 10.4 pp, low confidence
Gemini 1.5 Pro7.9%37.5%estimated ± 10.4 pp, low confidence
GLM-5.3-Flash41.8%90.3%estimated ± 10.4 pp, low confidence
GPT-4 Turbo7.0%34.4%estimated ± 10.4 pp, low confidence
GPT-4o mini6.7%33.0%estimated ± 10.4 pp, low confidence
o1-preview11.4%48.1%estimated ± 10.4 pp, low confidence
o3-pro21.9%69.5%estimated ± 10.4 pp, low confidence
Phi-4 Multimodal Instruct5.8%29.8%estimated ± 10.4 pp, low confidence
Qwen2.5 Coder 32B Instruct6.7%33.3%estimated ± 10.4 pp, low confidence