benchgap
Calibration

Artificial Analysis Intelligence Index → GPQA

GPQA is estimated from Artificial Analysis Intelligence Index with a Hill curve fitted on 62 models measured on both: y = 0.0000 + (0.9382 − 0.0000)·x^2.30 / (0.07365^2.30 + x^2.30), R² = 0.89, cross-validated error 3.2 pp. It is used for 90 estimates.

Estimated modelArtificial Analysis Intelligence IndexGPQASource
Apodex 1.1 Mini26.4%89.1%estimated ± 3.2 pp, high confidence
Celeris-16.4%39.0%estimated ± 3.2 pp, medium confidence
Claude 3 Haiku5.6%32.4%estimated ± 3.2 pp, medium confidence
Claude 3 Opus8.7%55.9%estimated ± 3.2 pp, high confidence
Claude 4.1 Opus18.6%83.8%estimated ± 3.2 pp, high confidence
Claude 4.1 Opus Thinking22.9%87.3%estimated ± 3.2 pp, high confidence
Claude 4 Sonnet16.6%81.3%estimated ± 3.2 pp, high confidence
Command A+22.5%87.1%estimated ± 3.2 pp, high confidence
DeepSeek-R113.1%74.2%estimated ± 3.2 pp, high confidence
DeepSeek R1 Distill Qwen 32B8.4%53.8%estimated ± 3.2 pp, high confidence
DeepSeek V3 03249.7%61.4%estimated ± 3.2 pp, high confidence
DeepSeek V3.113.7%75.7%estimated ± 3.2 pp, high confidence
DeepSeek V3.1 (Reasoning)13.5%75.1%estimated ± 3.2 pp, high confidence
DeepSeek V3.216.0%80.4%estimated ± 3.2 pp, high confidence
Exaone 4.0 1.2B5.2%29.3%estimated ± 3.2 pp, medium confidence
Exaone 4.0 32B6.3%38.8%estimated ± 3.2 pp, medium confidence
Gemini 1.0 Pro5.3%30.3%estimated ± 3.2 pp, medium confidence
Gemini 1.5 Pro7.9%50.7%estimated ± 3.2 pp, high confidence
Gemini 2.5 Flash9.9%62.0%estimated ± 3.2 pp, high confidence
Gemini 4 Argon52.6%92.8%estimated ± 3.2 pp, high confidence
Gemma 3 27B4.9%26.0%estimated ± 3.2 pp, medium confidence
GLM-4.5-Air11.1%67.5%estimated ± 3.2 pp, high confidence
GLM-5-Turbo26.6%89.2%estimated ± 3.2 pp, high confidence
GLM-5V-Turbo23.5%87.7%estimated ± 3.2 pp, high confidence
GPT-4 Turbo7.0%44.5%estimated ± 3.2 pp, medium confidence
GPT-4o8.4%54.2%estimated ± 3.2 pp, high confidence
GPT-4o mini6.7%41.5%estimated ± 3.2 pp, medium confidence
GPT-5.1-Codex23.7%87.8%estimated ± 3.2 pp, high confidence
GPT-5.1-Codex-Max23.7%87.8%estimated ± 3.2 pp, high confidence
GPT-5.2-Codex28.5%89.8%estimated ± 3.2 pp, high confidence
GPT-5.3 Codex32.5%90.8%estimated ± 3.2 pp, high confidence
GPT-5 (high)23.0%87.4%estimated ± 3.2 pp, high confidence
GPT-5 (medium)22.9%87.4%estimated ± 3.2 pp, high confidence
GPT-OSS 120B11.6%69.4%estimated ± 3.2 pp, high confidence
GPT-OSS 20B9.0%57.4%estimated ± 3.2 pp, high confidence
Granite-4.0-350M4.8%25.8%estimated ± 3.2 pp, medium confidence
Granite-4.0-H-1B5.2%29.0%estimated ± 3.2 pp, medium confidence
Granite-4.0-H-350M4.8%25.8%estimated ± 3.2 pp, medium confidence
Grok 422.5%87.1%estimated ± 3.2 pp, high confidence
Grok 4.1 Fast11.3%68.2%estimated ± 3.2 pp, high confidence
Grok 4.1 Fast (Reasoning)20.4%85.6%estimated ± 3.2 pp, high confidence
Grok 4 Fast (Reasoning)17.9%83.1%estimated ± 3.2 pp, high confidence
Grok Code Fast 114.1%76.5%estimated ± 3.2 pp, high confidence
Hy325.3%88.6%estimated ± 3.2 pp, high confidence
K-Exaone14.4%77.2%estimated ± 3.2 pp, high confidence
Kimi K212.7%73.0%estimated ± 3.2 pp, high confidence
Kimi K2.7 Code25.8%88.8%estimated ± 3.2 pp, high confidence
LFM2.5-2.6B8.4%53.9%estimated ± 3.2 pp, high confidence
LFM2.5-8B-A1B7.2%45.8%estimated ± 3.2 pp, medium confidence
LFM2.5-VL-1.6B-Extract4.8%25.8%estimated ± 3.2 pp, medium confidence
Ling 3.0 Flash VL24.6%88.3%estimated ± 3.2 pp, high confidence
Ling 3.0 Tiny11.1%67.4%estimated ± 3.2 pp, high confidence
Llama 3.1 405B7.3%46.4%estimated ± 3.2 pp, medium confidence
Llama 4 Maverick10.0%62.7%estimated ± 3.2 pp, high confidence
Llama 4 Scout8.1%51.9%estimated ± 3.2 pp, high confidence
MiMo-V2.6-Flash37.9%91.7%estimated ± 3.2 pp, high confidence
MiMo-V2.6-Pro46.3%92.5%estimated ± 3.2 pp, high confidence
MiMo-V2-Omni23.9%88.0%estimated ± 3.2 pp, high confidence
MiMo-V2-Pro28.6%89.9%estimated ± 3.2 pp, high confidence
Mistral Large 27.6%48.3%estimated ± 3.2 pp, medium confidence
Mistral Large 39.3%59.0%estimated ± 3.2 pp, high confidence
Mistral Large 438.4%91.8%estimated ± 3.2 pp, high confidence
Mistral Medium 39.0%57.8%estimated ± 3.2 pp, high confidence
Mistral Small 411.3%68.2%estimated ± 3.2 pp, high confidence
Mistral Small 4 (Reasoning)11.3%68.2%estimated ± 3.2 pp, high confidence
Muse Glimmer 30B17.5%82.5%estimated ± 3.2 pp, high confidence
Muse Spark 1.348.1%92.6%estimated ± 3.2 pp, high confidence
Nemotron 3 Nano 30B8.9%57.0%estimated ± 3.2 pp, high confidence
Nemotron 3 Super 100B12.8%73.3%estimated ± 3.2 pp, high confidence
Nemotron Ultra 253B7.5%48.1%estimated ± 3.2 pp, medium confidence
North Mini Code9.9%62.3%estimated ± 3.2 pp, high confidence
Nova Pro7.0%43.9%estimated ± 3.2 pp, medium confidence
o1-preview11.4%68.6%estimated ± 3.2 pp, high confidence
o320.2%85.4%estimated ± 3.2 pp, high confidence
o3-pro21.9%86.7%estimated ± 3.2 pp, high confidence
Phi-45.9%35.4%estimated ± 3.2 pp, medium confidence
Phi-4 Multimodal Instruct5.8%34.4%estimated ± 3.2 pp, medium confidence
Quasar 438B26.7%89.2%estimated ± 3.2 pp, high confidence
Qwen2.5 Coder 32B Instruct6.7%42.1%estimated ± 3.2 pp, medium confidence
Qwen3.5 397B (Reasoning)21.4%86.4%estimated ± 3.2 pp, high confidence
Qwen3.8 Max Preview45.4%92.4%estimated ± 3.2 pp, high confidence
Qwen3 Max15.6%79.7%estimated ± 3.2 pp, high confidence
Qwen3-Omni-30B-A3B-Instruct6.0%36.1%estimated ± 3.2 pp, medium confidence
Qwen3-Omni-30B-A3B-Thinking7.8%49.7%estimated ± 3.2 pp, medium confidence
Sarvam 105B8.8%56.3%estimated ± 3.2 pp, high confidence
Sarvam 30B6.6%40.7%estimated ± 3.2 pp, medium confidence
Solar Pro 27.0%44.2%estimated ± 3.2 pp, medium confidence
Solar Pro 37.8%50.1%estimated ± 3.2 pp, high confidence
Step 3.7 Flash19.5%84.8%estimated ± 3.2 pp, high confidence
Ultravox v0.6 Llama 3.3 70B7.7%49.0%estimated ± 3.2 pp, medium confidence