benchgap
Calibration

AA-GPQA Diamond → Vals GPQA Diamond

Vals GPQA Diamond is estimated from AA-GPQA Diamond with a inverse Michaelis–Menten curve fitted on 47 models measured on both: y = 1.08404·(x − 0.0000) / (0.0000 + 2.0000 − x), R² = 0.53, cross-validated error 7.4 pp. It is used for 86 estimates.

Estimated modelAA-GPQA DiamondVals GPQA DiamondSource
Apodex 1.186.4%82.4%estimated ± 7.4 pp, medium confidence
Apodex 1.1 Mini86.4%82.4%estimated ± 7.4 pp, medium confidence
Celeris-163.1%50.0%estimated ± 7.4 pp, low confidence
Claude 3 Haiku37.4%24.9%estimated ± 7.4 pp, low confidence
Claude 3 Opus48.9%35.1%estimated ± 7.4 pp, low confidence
Claude 4.1 Opus Thinking80.9%73.6%estimated ± 7.4 pp, medium confidence
Claude 4 Sonnet68.3%56.2%estimated ± 7.4 pp, medium confidence
Command A+76.1%66.6%estimated ± 7.4 pp, medium confidence
DeepSeek-R181.3%74.2%estimated ± 7.4 pp, medium confidence
DeepSeek R1 Distill Qwen 32B61.5%48.1%estimated ± 7.4 pp, low confidence
DeepSeek V3 032465.5%52.8%estimated ± 7.4 pp, medium confidence
DeepSeek V3.173.5%63.0%estimated ± 7.4 pp, medium confidence
DeepSeek V3.1 (Reasoning)77.9%69.2%estimated ± 7.4 pp, medium confidence
DeepSeek V3.275.1%65.2%estimated ± 7.4 pp, medium confidence
Exaone 4.0 1.2B42.4%29.2%estimated ± 7.4 pp, low confidence
Exaone 4.0 32B62.8%49.6%estimated ± 7.4 pp, low confidence
Gemini 1.0 Pro27.7%17.4%estimated ± 7.4 pp, low confidence
Gemini 1.5 Pro58.9%45.3%estimated ± 7.4 pp, low confidence
Gemini 2.5 Flash68.3%56.2%estimated ± 7.4 pp, medium confidence
Gemma 3 27B42.8%29.5%estimated ± 7.4 pp, low confidence
GLM-4.5-Air73.3%62.7%estimated ± 7.4 pp, medium confidence
GLM-5-Turbo84.7%79.6%estimated ± 7.4 pp, medium confidence
GLM-5V-Turbo80.9%73.6%estimated ± 7.4 pp, medium confidence
GPT-4o54.3%40.4%estimated ± 7.4 pp, low confidence
GPT-4o mini42.6%29.3%estimated ± 7.4 pp, low confidence
GPT-5.1-Codex86.0%81.8%estimated ± 7.4 pp, medium confidence
GPT-5.1-Codex-Max86.0%81.8%estimated ± 7.4 pp, medium confidence
GPT-5.2-Codex89.9%88.5%estimated ± 7.4 pp, medium confidence
GPT-5.3 Codex91.5%91.4%estimated ± 7.4 pp, medium confidence
GPT-5 (high)85.4%80.8%estimated ± 7.4 pp, medium confidence
GPT-5 (medium)84.2%78.8%estimated ± 7.4 pp, medium confidence
GPT-OSS 120B78.2%69.6%estimated ± 7.4 pp, medium confidence
GPT-OSS 20B68.8%56.8%estimated ± 7.4 pp, medium confidence
Granite-4.0-350M26.1%16.3%estimated ± 7.4 pp, low confidence
Granite-4.0-H-1B26.3%16.4%estimated ± 7.4 pp, low confidence
Granite-4.0-H-350M25.7%16.0%estimated ± 7.4 pp, low confidence
Grok 487.7%84.7%estimated ± 7.4 pp, medium confidence
Grok 4.1 Fast63.7%50.7%estimated ± 7.4 pp, medium confidence
Grok 4.1 Fast (Reasoning)85.3%80.6%estimated ± 7.4 pp, medium confidence
Grok 4 Fast (Reasoning)84.7%79.6%estimated ± 7.4 pp, medium confidence
Grok Code Fast 172.7%61.9%estimated ± 7.4 pp, medium confidence
Hy389.7%88.2%estimated ± 7.4 pp, medium confidence
K-Exaone78.3%69.7%estimated ± 7.4 pp, medium confidence
Kimi K276.6%67.3%estimated ± 7.4 pp, medium confidence
Kimi K2.7 Code89.6%88.0%estimated ± 7.4 pp, medium confidence
LFM2.5-2.6B55.8%41.9%estimated ± 7.4 pp, low confidence
LFM2.5-8B-A1B51.3%37.4%estimated ± 7.4 pp, low confidence
LFM2.5-VL-1.6B-Extract28.9%18.3%estimated ± 7.4 pp, low confidence
Ling 3.0 Flash VL86.2%82.1%estimated ± 7.4 pp, medium confidence
Ling 3.0 Tiny73.4%62.9%estimated ± 7.4 pp, medium confidence
Llama 3.1 405B51.5%37.6%estimated ± 7.4 pp, low confidence
Llama 4 Maverick67.1%54.7%estimated ± 7.4 pp, medium confidence
Llama 4 Scout58.7%45.0%estimated ± 7.4 pp, low confidence
MiMo-V2-Omni82.8%76.6%estimated ± 7.4 pp, medium confidence
MiMo-V2-Pro87.0%83.5%estimated ± 7.4 pp, medium confidence
Mistral Large 248.6%34.8%estimated ± 7.4 pp, low confidence
Mistral Large 368.0%55.8%estimated ± 7.4 pp, medium confidence
Mistral Medium 357.8%44.1%estimated ± 7.4 pp, low confidence
Mistral Small 476.9%67.7%estimated ± 7.4 pp, medium confidence
Mistral Small 4 (Reasoning)76.9%67.7%estimated ± 7.4 pp, medium confidence
Muse Glimmer 30B83.5%77.7%estimated ± 7.4 pp, medium confidence
Muse Spark 1.393.5%95.2%estimated ± 7.4 pp, medium confidence
Nemotron 3 Nano 30B75.7%66.0%estimated ± 7.4 pp, medium confidence
Nemotron 3 Super 100B80.0%72.3%estimated ± 7.4 pp, medium confidence
Nemotron Ultra 253B72.8%62.0%estimated ± 7.4 pp, medium confidence
North Mini Code75.7%66.0%estimated ± 7.4 pp, medium confidence
Nova Pro49.9%36.0%estimated ± 7.4 pp, low confidence
o1-preview76.5%67.1%estimated ± 7.4 pp, medium confidence
o382.7%76.4%estimated ± 7.4 pp, medium confidence
o3-pro84.5%79.3%estimated ± 7.4 pp, medium confidence
Phi-457.5%43.7%estimated ± 7.4 pp, low confidence
Phi-4 Multimodal Instruct31.5%20.3%estimated ± 7.4 pp, low confidence
Quasar 438B73.2%62.6%estimated ± 7.4 pp, medium confidence
Qwen2.5 Coder 32B Instruct41.7%28.6%estimated ± 7.4 pp, low confidence
Qwen3.5 397B (Reasoning)86.1%81.9%estimated ± 7.4 pp, medium confidence
Qwen 3.6 Max (preview)88.8%86.6%estimated ± 7.4 pp, medium confidence
Qwen3.8 Max Preview92.8%93.8%estimated ± 7.4 pp, medium confidence
Qwen3 Max76.4%67.0%estimated ± 7.4 pp, medium confidence
Qwen3-Omni-30B-A3B-Instruct62.0%48.7%estimated ± 7.4 pp, low confidence
Qwen3-Omni-30B-A3B-Thinking72.6%61.8%estimated ± 7.4 pp, medium confidence
Sarvam 105B73.8%63.4%estimated ± 7.4 pp, medium confidence
Sarvam 30B63.3%50.2%estimated ± 7.4 pp, medium confidence
Solar Pro 256.1%42.3%estimated ± 7.4 pp, low confidence
Solar Pro 372.4%61.5%estimated ± 7.4 pp, medium confidence
Step 3.7 Flash80.9%73.6%estimated ± 7.4 pp, medium confidence
Ultravox v0.6 Llama 3.3 70B49.8%35.9%estimated ± 7.4 pp, low confidence