benchgap
Calibration

AA-HLE → GPQA Diamond

GPQA Diamond is estimated from AA-HLE with a Michaelis–Menten + offset curve fitted on 43 models measured on both: y = 0.6441 + 0.6612·x / (0.58299 + x), R² = 0.82, cross-validated error 2.2 pp. It is used for 89 estimates.

Estimated modelAA-HLEGPQA DiamondSource
Apodex 1.134.1%88.8%estimated ± 2.2 pp, high confidence
Apodex 1.1 Mini34.1%88.8%estimated ± 2.2 pp, high confidence
Celeris-16.8%71.3%estimated ± 2.2 pp, high confidence
Claude 3 Haiku4.1%68.8%estimated ± 2.2 pp, medium confidence
Claude 3 Opus2.8%67.4%estimated ± 2.2 pp, medium confidence
Claude 4.1 Opus Thinking12.5%76.1%estimated ± 2.2 pp, high confidence
Claude 4 Sonnet4.3%69.0%estimated ± 2.2 pp, medium confidence
Command A+12.0%75.7%estimated ± 2.2 pp, high confidence
DeepSeek-R115.8%78.5%estimated ± 2.2 pp, high confidence
DeepSeek R1 Distill Qwen 32B4.6%69.2%estimated ± 2.2 pp, medium confidence
DeepSeek V3 03244.7%69.3%estimated ± 2.2 pp, medium confidence
DeepSeek V3.16.7%71.2%estimated ± 2.2 pp, high confidence
DeepSeek V3.1 (Reasoning)14.3%77.4%estimated ± 2.2 pp, high confidence
DeepSeek V3.211.2%75.1%estimated ± 2.2 pp, high confidence
Exaone 4.0 1.2B5.7%70.3%estimated ± 2.2 pp, high confidence
Exaone 4.0 32B5.0%69.6%estimated ± 2.2 pp, high confidence
Gemini 1.0 Pro4.2%68.9%estimated ± 2.2 pp, medium confidence
Gemini 1.5 Pro4.6%69.2%estimated ± 2.2 pp, medium confidence
Gemini 2.5 Flash4.7%69.3%estimated ± 2.2 pp, medium confidence
Gemini 4 Argon57.1%97.1%estimated ± 2.2 pp, medium confidence
Gemma 3 27B4.4%69.0%estimated ± 2.2 pp, medium confidence
GLM-4.5-Air7.0%71.5%estimated ± 2.2 pp, high confidence
GLM-5-Turbo27.8%85.8%estimated ± 2.2 pp, high confidence
GLM-5V-Turbo17.1%79.4%estimated ± 2.2 pp, high confidence
GPT-4 Turbo3.1%67.7%estimated ± 2.2 pp, medium confidence
GPT-4o2.4%67.0%estimated ± 2.2 pp, medium confidence
GPT-4o mini4.2%68.9%estimated ± 2.2 pp, medium confidence
GPT-5.1-Codex25.7%84.6%estimated ± 2.2 pp, high confidence
GPT-5.1-Codex-Max25.7%84.6%estimated ± 2.2 pp, high confidence
GPT-5.2-Codex35.7%89.5%estimated ± 2.2 pp, high confidence
GPT-5.3 Codex42.5%92.3%estimated ± 2.2 pp, high confidence
GPT-5 (high)28.5%86.1%estimated ± 2.2 pp, high confidence
GPT-5 (medium)25.4%84.5%estimated ± 2.2 pp, high confidence
GPT-OSS 120B19.6%81.0%estimated ± 2.2 pp, high confidence
GPT-OSS 20B11.0%74.9%estimated ± 2.2 pp, high confidence
Granite-4.0-350M5.5%70.1%estimated ± 2.2 pp, high confidence
Granite-4.0-H-1B5.0%69.6%estimated ± 2.2 pp, high confidence
Granite-4.0-H-350M6.4%71.0%estimated ± 2.2 pp, high confidence
Grok 426.7%85.2%estimated ± 2.2 pp, high confidence
Grok 4.1 Fast5.1%69.7%estimated ± 2.2 pp, high confidence
Grok 4.1 Fast (Reasoning)19.3%80.9%estimated ± 2.2 pp, high confidence
Grok 4 Fast (Reasoning)19.1%80.7%estimated ± 2.2 pp, high confidence
Grok Code Fast 18.0%72.4%estimated ± 2.2 pp, high confidence
Hy333.5%88.5%estimated ± 2.2 pp, high confidence
K-Exaone13.9%77.1%estimated ± 2.2 pp, high confidence
Kimi K27.4%71.9%estimated ± 2.2 pp, high confidence
Kimi K2.7 Code35.0%89.2%estimated ± 2.2 pp, high confidence
LFM2.5-2.6B6.2%70.8%estimated ± 2.2 pp, high confidence
LFM2.5-8B-A1B6.9%71.4%estimated ± 2.2 pp, high confidence
LFM2.5-VL-1.6B-Extract5.1%69.7%estimated ± 2.2 pp, high confidence
Ling 3.0 Flash VL22.0%82.5%estimated ± 2.2 pp, high confidence
Ling 3.0 Tiny9.3%73.5%estimated ± 2.2 pp, high confidence
Llama 3.1 405B4.0%68.7%estimated ± 2.2 pp, medium confidence
Llama 4 Maverick4.9%69.5%estimated ± 2.2 pp, high confidence
Llama 4 Scout3.8%68.5%estimated ± 2.2 pp, medium confidence
MiMo-V2.6-Flash35.1%89.3%estimated ± 2.2 pp, high confidence
MiMo-V2.6-Pro49.4%94.7%estimated ± 2.2 pp, high confidence
MiMo-V2-Omni22.1%82.6%estimated ± 2.2 pp, high confidence
MiMo-V2-Pro30.4%87.1%estimated ± 2.2 pp, high confidence
Mistral Large 23.3%68.0%estimated ± 2.2 pp, medium confidence
Mistral Large 34.2%68.9%estimated ± 2.2 pp, medium confidence
Mistral Large 435.0%89.2%estimated ± 2.2 pp, high confidence
Mistral Medium 34.1%68.8%estimated ± 2.2 pp, medium confidence
Mistral Small 49.9%74.0%estimated ± 2.2 pp, high confidence
Mistral Small 4 (Reasoning)9.9%74.0%estimated ± 2.2 pp, high confidence
Muse Glimmer 30B22.0%82.5%estimated ± 2.2 pp, high confidence
Muse Spark 1.348.7%94.5%estimated ± 2.2 pp, high confidence
Nemotron 3 Nano 30B11.4%75.2%estimated ± 2.2 pp, high confidence
Nemotron 3 Super 100B20.8%81.8%estimated ± 2.2 pp, high confidence
Nemotron Ultra 253B7.4%71.9%estimated ± 2.2 pp, high confidence
North Mini Code11.1%75.0%estimated ± 2.2 pp, high confidence
Nova Pro3.2%67.9%estimated ± 2.2 pp, medium confidence
o320.1%81.4%estimated ± 2.2 pp, high confidence
Phi-43.8%68.5%estimated ± 2.2 pp, medium confidence
Phi-4 Multimodal Instruct5.0%69.6%estimated ± 2.2 pp, high confidence
Quasar 438B18.7%80.5%estimated ± 2.2 pp, high confidence
Qwen2.5 Coder 32B Instruct3.5%68.2%estimated ± 2.2 pp, medium confidence
Qwen3.5 397B (Reasoning)19.8%81.2%estimated ± 2.2 pp, high confidence
Qwen 3.6 Max (preview)30.8%87.3%estimated ± 2.2 pp, high confidence
Qwen3.8 Max Preview43.1%92.5%estimated ± 2.2 pp, high confidence
Qwen3 Max11.9%75.6%estimated ± 2.2 pp, high confidence
Qwen3-Omni-30B-A3B-Instruct4.6%69.2%estimated ± 2.2 pp, medium confidence
Qwen3-Omni-30B-A3B-Thinking7.5%71.9%estimated ± 2.2 pp, high confidence
Sarvam 105B11.0%74.9%estimated ± 2.2 pp, high confidence
Sarvam 30B7.5%71.9%estimated ± 2.2 pp, high confidence
Solar Pro 23.7%68.4%estimated ± 2.2 pp, medium confidence
Solar Pro 310.3%74.3%estimated ± 2.2 pp, high confidence
Step 3.7 Flash21.4%82.2%estimated ± 2.2 pp, high confidence
Ultravox v0.6 Llama 3.3 70B3.6%68.3%estimated ± 2.2 pp, medium confidence