benchgap
Calibration

AA-GPQA Diamond → HLE-Verified

HLE-Verified is estimated from AA-GPQA Diamond with a offset logistic curve fitted on 6 models measured on both: y = 0.1532 + (0.5458 − 0.1532) / (1 + exp(−200.00·(x − 0.9130))), R² = 1.00, cross-validated error 2.0 pp. It is used for 150 estimates.

Estimated modelAA-GPQA DiamondHLE-VerifiedSource
A.X K285.7%15.3%estimated ± 2.0 pp, low confidence
Apodex 1.186.4%15.3%estimated ± 2.0 pp, low confidence
Apodex 1.1 Mini86.4%15.3%estimated ± 2.0 pp, low confidence
Celeris-163.1%15.3%estimated ± 2.0 pp, low confidence
Claude 3 Haiku37.4%15.3%estimated ± 2.0 pp, low confidence
Claude 3 Opus48.9%15.3%estimated ± 2.0 pp, low confidence
Claude 4.1 Opus Thinking80.9%15.3%estimated ± 2.0 pp, low confidence
Claude 4 Sonnet68.3%15.3%estimated ± 2.0 pp, low confidence
Claude Opus 4.581.0%15.3%estimated ± 2.0 pp, low confidence
Claude Opus 4.684.0%15.3%estimated ± 2.0 pp, low confidence
Claude Opus 4.788.5%15.5%estimated ± 2.0 pp, low confidence
Command A+76.1%15.3%estimated ± 2.0 pp, low confidence
DeepSeek-R181.3%15.3%estimated ± 2.0 pp, low confidence
DeepSeek R1 Distill Qwen 32B61.5%15.3%estimated ± 2.0 pp, low confidence
DeepSeek V355.7%15.3%estimated ± 2.0 pp, low confidence
DeepSeek V3 032465.5%15.3%estimated ± 2.0 pp, low confidence
DeepSeek V3.173.5%15.3%estimated ± 2.0 pp, low confidence
DeepSeek V3.1 (Reasoning)77.9%15.3%estimated ± 2.0 pp, low confidence
DeepSeek V3.275.1%15.3%estimated ± 2.0 pp, low confidence
Exaone 4.0 1.2B42.4%15.3%estimated ± 2.0 pp, low confidence
Exaone 4.0 32B62.8%15.3%estimated ± 2.0 pp, low confidence
Gemini 1.0 Pro27.7%15.3%estimated ± 2.0 pp, low confidence
Gemini 1.5 Pro58.9%15.3%estimated ± 2.0 pp, low confidence
Gemini 2.5 Flash68.3%15.3%estimated ± 2.0 pp, low confidence
Gemini 2.5 Pro84.4%15.3%estimated ± 2.0 pp, low confidence
Gemini 3.5 Flash-Lite83.8%15.3%estimated ± 2.0 pp, low confidence
Gemini 3 Flash81.2%15.3%estimated ± 2.0 pp, low confidence
Gemma 3 27B42.8%15.3%estimated ± 2.0 pp, low confidence
Gemma 4 12B75.3%15.3%estimated ± 2.0 pp, low confidence
Gemma 4 26B A4B79.2%15.3%estimated ± 2.0 pp, low confidence
Gemma 4 31B85.7%15.3%estimated ± 2.0 pp, low confidence
Gemma 4 E2B43.3%15.3%estimated ± 2.0 pp, low confidence
Gemma 4 E4B57.6%15.3%estimated ± 2.0 pp, low confidence
GLM-4.5-Air73.3%15.3%estimated ± 2.0 pp, low confidence
GLM-4.663.2%15.3%estimated ± 2.0 pp, low confidence
GLM-4.785.9%15.3%estimated ± 2.0 pp, low confidence
GLM-582.0%15.3%estimated ± 2.0 pp, low confidence
GLM-5.186.8%15.3%estimated ± 2.0 pp, low confidence
GLM-5.289.5%16.4%estimated ± 2.0 pp, low confidence
GLM-5.391.7%42.3%estimated ± 2.0 pp, medium confidence
GLM-5.3-Flash91.2%32.9%estimated ± 2.0 pp, medium confidence
GLM-5-Turbo84.7%15.3%estimated ± 2.0 pp, low confidence
GLM-5V-Turbo80.9%15.3%estimated ± 2.0 pp, low confidence
GPT-4.166.6%15.3%estimated ± 2.0 pp, low confidence
GPT-4.1 mini66.4%15.3%estimated ± 2.0 pp, low confidence
GPT-4.1 nano51.2%15.3%estimated ± 2.0 pp, low confidence
GPT-4o54.3%15.3%estimated ± 2.0 pp, low confidence
GPT-4o mini42.6%15.3%estimated ± 2.0 pp, low confidence
GPT-5.187.3%15.3%estimated ± 2.0 pp, low confidence
GPT-5.1-Codex86.0%15.3%estimated ± 2.0 pp, low confidence
GPT-5.1-Codex-Max86.0%15.3%estimated ± 2.0 pp, low confidence
GPT-5.2-Codex89.9%17.6%estimated ± 2.0 pp, low confidence
GPT-5.3 Codex91.5%38.7%estimated ± 2.0 pp, medium confidence
GPT-5 (high)85.4%15.3%estimated ± 2.0 pp, low confidence
GPT-5 (medium)84.2%15.3%estimated ± 2.0 pp, low confidence
GPT-OSS 120B78.2%15.3%estimated ± 2.0 pp, low confidence
GPT-OSS 20B68.8%15.3%estimated ± 2.0 pp, low confidence
Granite-4.0-350M26.1%15.3%estimated ± 2.0 pp, low confidence
Granite-4.0-H-1B26.3%15.3%estimated ± 2.0 pp, low confidence
Granite-4.0-H-350M25.7%15.3%estimated ± 2.0 pp, low confidence
Granite 4.2 30B64.4%15.3%estimated ± 2.0 pp, low confidence
Granite 4.2 3B55.9%15.3%estimated ± 2.0 pp, low confidence
Granite 4.2 8B63.1%15.3%estimated ± 2.0 pp, low confidence
Grok 487.7%15.3%estimated ± 2.0 pp, low confidence
Grok 4.1 Fast63.7%15.3%estimated ± 2.0 pp, low confidence
Grok 4.1 Fast (Reasoning)85.3%15.3%estimated ± 2.0 pp, low confidence
Grok 4.390.1%18.6%estimated ± 2.0 pp, low confidence
Grok 4 Fast (Reasoning)84.7%15.3%estimated ± 2.0 pp, low confidence
Grok Code Fast 172.7%15.3%estimated ± 2.0 pp, low confidence
Hy389.7%16.8%estimated ± 2.0 pp, low confidence
Hy3 Preview89.7%16.8%estimated ± 2.0 pp, low confidence
Inkling87.2%15.3%estimated ± 2.0 pp, low confidence
K-Exaone78.3%15.3%estimated ± 2.0 pp, low confidence
K-EXAONE 2.082.9%15.3%estimated ± 2.0 pp, low confidence
Kimi K2.691.1%31.0%estimated ± 2.0 pp, medium confidence
Kimi K276.6%15.3%estimated ± 2.0 pp, low confidence
Kimi K2.587.9%15.4%estimated ± 2.0 pp, low confidence
Kimi K2.5 (Reasoning)87.9%15.4%estimated ± 2.0 pp, low confidence
Kimi K2.7 Code89.6%16.6%estimated ± 2.0 pp, low confidence
LFM2.5-2.6B55.8%15.3%estimated ± 2.0 pp, low confidence
LFM2.5-8B-A1B51.3%15.3%estimated ± 2.0 pp, low confidence
LFM2.5-VL-1.6B-Extract28.9%15.3%estimated ± 2.0 pp, low confidence
Ling 2.6 Flash59.3%15.3%estimated ± 2.0 pp, low confidence
Ling 3.0 Flash85.5%15.3%estimated ± 2.0 pp, low confidence
Ling 3.0 Flash FP885.5%15.3%estimated ± 2.0 pp, low confidence
Ling 3.0 Flash VL86.2%15.3%estimated ± 2.0 pp, low confidence
Ling 3.0 Tiny73.4%15.3%estimated ± 2.0 pp, low confidence
Llama 3.1 405B51.5%15.3%estimated ± 2.0 pp, low confidence
Llama 4 Maverick67.1%15.3%estimated ± 2.0 pp, low confidence
Llama 4 Scout58.7%15.3%estimated ± 2.0 pp, low confidence
MiMo-V2.5-Pro86.6%15.3%estimated ± 2.0 pp, low confidence
MiMo-V2-Flash65.6%15.3%estimated ± 2.0 pp, low confidence
MiMo-V2-Omni82.8%15.3%estimated ± 2.0 pp, low confidence
MiMo-V2-Pro87.0%15.3%estimated ± 2.0 pp, low confidence
MiniCPM5-2B70.2%15.3%estimated ± 2.0 pp, low confidence
MiniMax M2.787.4%15.3%estimated ± 2.0 pp, low confidence
MiniMax M392.9%53.0%estimated ± 2.0 pp, medium confidence
Mistral Large 248.6%15.3%estimated ± 2.0 pp, low confidence
Mistral Large 368.0%15.3%estimated ± 2.0 pp, low confidence
Mistral Medium 357.8%15.3%estimated ± 2.0 pp, low confidence
Mistral Medium 3.5 128B74.8%15.3%estimated ± 2.0 pp, low confidence
Mistral Small 476.9%15.3%estimated ± 2.0 pp, low confidence
Mistral Small 4 (Reasoning)76.9%15.3%estimated ± 2.0 pp, low confidence
Muse Glimmer 30B83.5%15.3%estimated ± 2.0 pp, low confidence
Muse Spark 1.189.8%17.2%estimated ± 2.0 pp, low confidence
Muse Spark 1.290.4%20.8%estimated ± 2.0 pp, low confidence
Muse Spark 1.393.5%54.1%estimated ± 2.0 pp, medium confidence
Nemotron 3.5 Lightning 30B A3B NVFP474.3%15.3%estimated ± 2.0 pp, low confidence
Nemotron 3 Nano 30B75.7%15.3%estimated ± 2.0 pp, low confidence
Nemotron 3 Nano Omni 30B A3B46.9%15.3%estimated ± 2.0 pp, low confidence
Nemotron 3 Super 100B80.0%15.3%estimated ± 2.0 pp, low confidence
Nemotron 3 Ultra86.7%15.3%estimated ± 2.0 pp, low confidence
Nemotron Ultra 253B72.8%15.3%estimated ± 2.0 pp, low confidence
North Mini Code75.7%15.3%estimated ± 2.0 pp, low confidence
Nova Pro49.9%15.3%estimated ± 2.0 pp, low confidence
o174.7%15.3%estimated ± 2.0 pp, low confidence
o1-preview76.5%15.3%estimated ± 2.0 pp, low confidence
o382.7%15.3%estimated ± 2.0 pp, low confidence
o3-mini74.8%15.3%estimated ± 2.0 pp, low confidence
o3-pro84.5%15.3%estimated ± 2.0 pp, low confidence
Phi-457.5%15.3%estimated ± 2.0 pp, low confidence
Phi-4 Multimodal Instruct31.5%15.3%estimated ± 2.0 pp, low confidence
Quasar 438B73.2%15.3%estimated ± 2.0 pp, low confidence
Qwen2.5 Coder 32B Instruct41.7%15.3%estimated ± 2.0 pp, low confidence
Qwen3.5-122B-A10B85.7%15.3%estimated ± 2.0 pp, low confidence
Qwen3.5-27B85.8%15.3%estimated ± 2.0 pp, low confidence
Qwen3.5-35B-A3B84.5%15.3%estimated ± 2.0 pp, low confidence
Qwen3.5 397B86.1%15.3%estimated ± 2.0 pp, low confidence
Qwen3.5 397B (Reasoning)86.1%15.3%estimated ± 2.0 pp, low confidence
Qwen3.6-27B84.2%15.3%estimated ± 2.0 pp, low confidence
Qwen3.6-35B-A3B84.1%15.3%estimated ± 2.0 pp, low confidence
Qwen 3.6 Max (preview)88.8%15.6%estimated ± 2.0 pp, low confidence
Qwen3.6 Plus88.2%15.4%estimated ± 2.0 pp, low confidence
Qwen3.7 Max92.3%49.9%estimated ± 2.0 pp, medium confidence
Qwen3.7 Plus90.0%18.0%estimated ± 2.0 pp, low confidence
Qwen3.8-27B90.5%21.9%estimated ± 2.0 pp, low confidence
Qwen3.8-Flash-Next92.3%49.9%estimated ± 2.0 pp, medium confidence
Qwen3.8 Max Preview92.8%52.7%estimated ± 2.0 pp, medium confidence
Qwen3 Max76.4%15.3%estimated ± 2.0 pp, low confidence
Qwen3-Omni-30B-A3B-Instruct62.0%15.3%estimated ± 2.0 pp, low confidence
Qwen3-Omni-30B-A3B-Thinking72.6%15.3%estimated ± 2.0 pp, low confidence
Sarvam 105B73.8%15.3%estimated ± 2.0 pp, low confidence
Sarvam 30B63.3%15.3%estimated ± 2.0 pp, low confidence
Solar Pro 256.1%15.3%estimated ± 2.0 pp, low confidence
Solar Pro 372.4%15.3%estimated ± 2.0 pp, low confidence
Solar Pro 489.1%15.8%estimated ± 2.0 pp, low confidence
Step 3.7 Flash80.9%15.3%estimated ± 2.0 pp, low confidence
Trinity-Large-Preview75.2%15.3%estimated ± 2.0 pp, low confidence
Trinity-Large-Thinking75.2%15.3%estimated ± 2.0 pp, low confidence
Ultravox v0.6 Llama 3.3 70B49.8%15.3%estimated ± 2.0 pp, low confidence