benchgap
Calibration

AA-GPQA Diamond → LABBench2

LABBench2 is estimated from AA-GPQA Diamond with a inverse Michaelis–Menten curve fitted on 6 models measured on both: y = 0.94179·(x − 0.0000) / (0.0000 + 2.0000 − x), R² = 0.44, cross-validated error 1.8 pp. It is used for 127 estimates.

Estimated modelAA-GPQA DiamondLABBench2Source
A.X K285.7%70.6%estimated ± 1.8 pp, low confidence
Apodex 1.186.4%71.6%estimated ± 1.8 pp, low confidence
Apodex 1.1 Mini86.4%71.6%estimated ± 1.8 pp, low confidence
Celeris-163.1%43.4%estimated ± 1.8 pp, low confidence
Claude 3 Haiku37.4%21.7%estimated ± 1.8 pp, low confidence
Claude 3 Opus48.9%30.5%estimated ± 1.8 pp, low confidence
Claude 4.1 Opus Thinking80.9%64.0%estimated ± 1.8 pp, low confidence
Claude 4 Sonnet68.3%48.8%estimated ± 1.8 pp, low confidence
Claude Opus 4.581.0%64.1%estimated ± 1.8 pp, low confidence
Claude Opus 4.684.0%68.2%estimated ± 1.8 pp, low confidence
Claude Opus 4.7 (Adaptive)91.4%79.3%estimated ± 1.8 pp, low confidence
Command A+76.1%57.8%estimated ± 1.8 pp, low confidence
DeepSeek-R181.3%64.5%estimated ± 1.8 pp, low confidence
DeepSeek R1 Distill Qwen 32B61.5%41.8%estimated ± 1.8 pp, low confidence
DeepSeek V355.7%36.4%estimated ± 1.8 pp, low confidence
DeepSeek V3 032465.5%45.9%estimated ± 1.8 pp, low confidence
DeepSeek V3.173.5%54.7%estimated ± 1.8 pp, low confidence
DeepSeek V3.1 (Reasoning)77.9%60.1%estimated ± 1.8 pp, low confidence
DeepSeek V3.275.1%56.6%estimated ± 1.8 pp, low confidence
Exaone 4.0 1.2B42.4%25.3%estimated ± 1.8 pp, low confidence
Exaone 4.0 32B62.8%43.1%estimated ± 1.8 pp, low confidence
Gemini 1.0 Pro27.7%15.1%estimated ± 1.8 pp, low confidence
Gemini 1.5 Pro58.9%39.3%estimated ± 1.8 pp, low confidence
Gemini 2.5 Flash68.3%48.8%estimated ± 1.8 pp, low confidence
Gemini 2.5 Pro84.4%68.8%estimated ± 1.8 pp, low confidence
Gemma 3 27B42.8%25.6%estimated ± 1.8 pp, low confidence
Gemma 4 12B75.3%56.9%estimated ± 1.8 pp, low confidence
Gemma 4 26B A4B79.2%61.7%estimated ± 1.8 pp, low confidence
Gemma 4 31B85.7%70.6%estimated ± 1.8 pp, low confidence
Gemma 4 E2B43.3%26.0%estimated ± 1.8 pp, low confidence
Gemma 4 E4B57.6%38.1%estimated ± 1.8 pp, low confidence
GLM-4.5-Air73.3%54.5%estimated ± 1.8 pp, low confidence
GLM-582.0%65.4%estimated ± 1.8 pp, low confidence
GLM-5-Turbo84.7%69.2%estimated ± 1.8 pp, low confidence
GLM-5V-Turbo80.9%64.0%estimated ± 1.8 pp, low confidence
GPT-4.166.6%47.0%estimated ± 1.8 pp, low confidence
GPT-4.1 mini66.4%46.8%estimated ± 1.8 pp, low confidence
GPT-4.1 nano51.2%32.4%estimated ± 1.8 pp, low confidence
GPT-4o54.3%35.1%estimated ± 1.8 pp, low confidence
GPT-4o mini42.6%25.5%estimated ± 1.8 pp, low confidence
GPT-5.1-Codex86.0%71.0%estimated ± 1.8 pp, low confidence
GPT-5.1-Codex-Max86.0%71.0%estimated ± 1.8 pp, low confidence
GPT-5.2-Codex89.9%76.9%estimated ± 1.8 pp, low confidence
GPT-5.3 Codex91.5%79.4%estimated ± 1.8 pp, low confidence
GPT-5 (high)85.4%70.2%estimated ± 1.8 pp, low confidence
GPT-5 (medium)84.2%68.5%estimated ± 1.8 pp, low confidence
GPT-OSS 120B78.2%60.5%estimated ± 1.8 pp, low confidence
GPT-OSS 20B68.8%49.4%estimated ± 1.8 pp, low confidence
Granite-4.0-350M26.1%14.1%estimated ± 1.8 pp, low confidence
Granite-4.0-H-1B26.3%14.3%estimated ± 1.8 pp, low confidence
Granite-4.0-H-350M25.7%13.9%estimated ± 1.8 pp, low confidence
Granite 4.2 30B64.4%44.7%estimated ± 1.8 pp, low confidence
Granite 4.2 3B55.9%36.5%estimated ± 1.8 pp, low confidence
Granite 4.2 8B63.1%43.4%estimated ± 1.8 pp, low confidence
Grok 487.7%73.5%estimated ± 1.8 pp, low confidence
Grok 4.1 Fast63.7%44.0%estimated ± 1.8 pp, low confidence
Grok 4.1 Fast (Reasoning)85.3%70.0%estimated ± 1.8 pp, low confidence
Grok 4 Fast (Reasoning)84.7%69.2%estimated ± 1.8 pp, low confidence
Grok Code Fast 172.7%53.8%estimated ± 1.8 pp, low confidence
Hy389.7%76.6%estimated ± 1.8 pp, low confidence
Hy3 Preview89.7%76.6%estimated ± 1.8 pp, low confidence
K-Exaone78.3%60.6%estimated ± 1.8 pp, low confidence
K-EXAONE 2.082.9%66.7%estimated ± 1.8 pp, low confidence
Kimi K276.6%58.5%estimated ± 1.8 pp, low confidence
Kimi K2.587.9%73.8%estimated ± 1.8 pp, low confidence
Kimi K2.5 (Reasoning)87.9%73.8%estimated ± 1.8 pp, low confidence
Kimi K2.7 Code89.6%76.4%estimated ± 1.8 pp, low confidence
LFM2.5-2.6B55.8%36.4%estimated ± 1.8 pp, low confidence
LFM2.5-8B-A1B51.3%32.5%estimated ± 1.8 pp, low confidence
LFM2.5-VL-1.6B-Extract28.9%15.9%estimated ± 1.8 pp, low confidence
Ling 2.6 Flash59.3%39.7%estimated ± 1.8 pp, low confidence
Ling 3.0 Flash FP885.5%70.3%estimated ± 1.8 pp, low confidence
Ling 3.0 Flash VL86.2%71.3%estimated ± 1.8 pp, low confidence
Ling 3.0 Tiny73.4%54.6%estimated ± 1.8 pp, low confidence
Llama 3.1 405B51.5%32.7%estimated ± 1.8 pp, low confidence
Llama 4 Maverick67.1%47.5%estimated ± 1.8 pp, low confidence
Llama 4 Scout58.7%39.1%estimated ± 1.8 pp, low confidence
MiMo-V2-Flash65.6%46.0%estimated ± 1.8 pp, low confidence
MiMo-V2-Omni82.8%66.5%estimated ± 1.8 pp, low confidence
MiMo-V2-Pro87.0%72.5%estimated ± 1.8 pp, low confidence
MiniCPM5-2B70.2%50.9%estimated ± 1.8 pp, low confidence
Mistral Large 248.6%30.2%estimated ± 1.8 pp, low confidence
Mistral Large 368.0%48.5%estimated ± 1.8 pp, low confidence
Mistral Medium 357.8%38.3%estimated ± 1.8 pp, low confidence
Mistral Small 476.9%58.8%estimated ± 1.8 pp, low confidence
Mistral Small 4 (Reasoning)76.9%58.8%estimated ± 1.8 pp, low confidence
Muse Glimmer 30B83.5%67.5%estimated ± 1.8 pp, low confidence
Muse Spark 1.393.5%82.7%estimated ± 1.8 pp, low confidence
Nemotron 3.5 Lightning 30B A3B NVFP474.3%55.7%estimated ± 1.8 pp, low confidence
Nemotron 3 Nano 30B75.7%57.4%estimated ± 1.8 pp, low confidence
Nemotron 3 Nano Omni 30B A3B46.9%28.9%estimated ± 1.8 pp, low confidence
Nemotron 3 Super 100B80.0%62.8%estimated ± 1.8 pp, low confidence
Nemotron Ultra 253B72.8%53.9%estimated ± 1.8 pp, low confidence
North Mini Code75.7%57.4%estimated ± 1.8 pp, low confidence
Nova Pro49.9%31.3%estimated ± 1.8 pp, low confidence
o174.7%56.1%estimated ± 1.8 pp, low confidence
o1-preview76.5%58.3%estimated ± 1.8 pp, low confidence
o382.7%66.4%estimated ± 1.8 pp, low confidence
o3-mini74.8%56.3%estimated ± 1.8 pp, low confidence
o3-pro84.5%68.9%estimated ± 1.8 pp, low confidence
Phi-457.5%38.0%estimated ± 1.8 pp, low confidence
Phi-4 Multimodal Instruct31.5%17.6%estimated ± 1.8 pp, low confidence
Quasar 438B73.2%54.4%estimated ± 1.8 pp, low confidence
Qwen2.5 Coder 32B Instruct41.7%24.8%estimated ± 1.8 pp, low confidence
Qwen3.5-122B-A10B85.7%70.6%estimated ± 1.8 pp, low confidence
Qwen3.5-27B85.8%70.8%estimated ± 1.8 pp, low confidence
Qwen3.5-35B-A3B84.5%68.9%estimated ± 1.8 pp, low confidence
Qwen3.5 397B86.1%71.2%estimated ± 1.8 pp, low confidence
Qwen3.5 397B (Reasoning)86.1%71.2%estimated ± 1.8 pp, low confidence
Qwen3.6-27B84.2%68.5%estimated ± 1.8 pp, low confidence
Qwen3.6-35B-A3B84.1%68.3%estimated ± 1.8 pp, low confidence
Qwen 3.6 Max (preview)88.8%75.2%estimated ± 1.8 pp, low confidence
Qwen3.7 Plus90.0%77.1%estimated ± 1.8 pp, low confidence
Qwen3.8-Flash-Next92.3%80.7%estimated ± 1.8 pp, low confidence
Qwen3.8 Max Preview92.8%81.5%estimated ± 1.8 pp, low confidence
Qwen3 Max76.4%58.2%estimated ± 1.8 pp, low confidence
Qwen3-Omni-30B-A3B-Instruct62.0%42.3%estimated ± 1.8 pp, low confidence
Qwen3-Omni-30B-A3B-Thinking72.6%53.7%estimated ± 1.8 pp, low confidence
Sarvam 105B73.8%55.1%estimated ± 1.8 pp, low confidence
Sarvam 30B63.3%43.6%estimated ± 1.8 pp, low confidence
Solar Pro 256.1%36.7%estimated ± 1.8 pp, low confidence
Solar Pro 372.4%53.4%estimated ± 1.8 pp, low confidence
Solar Pro 489.1%75.7%estimated ± 1.8 pp, low confidence
Step 3.7 Flash80.9%64.0%estimated ± 1.8 pp, low confidence
Trinity-Large-Preview75.2%56.7%estimated ± 1.8 pp, low confidence
Trinity-Large-Thinking75.2%56.7%estimated ± 1.8 pp, low confidence
Ultravox v0.6 Llama 3.3 70B49.8%31.2%estimated ± 1.8 pp, low confidence