benchgap
Calibration

GDPval-AA → AA-AnalystAgent

AA-AnalystAgent is estimated from GDPval-AA with a inverse Michaelis–Menten curve fitted on 15 models measured on both: y = 1.22077·(x − 0.0000) / (0.0000 + 2.0000 − x), R² = 0.55, cross-validated error 12.0 pp. It is used for 75 estimates.

Estimated modelGDPval-AAAA-AnalystAgentSource
A.X K223.2%16.0%estimated ± 12.0 pp, low confidence
Apodex 1.1 Mini34.8%25.7%estimated ± 12.0 pp, low confidence
Celeris-10.0%0.0%estimated ± 12.0 pp, low confidence
Command A+0.9%0.6%estimated ± 12.0 pp, low confidence
DeepSeek V30.0%0.0%estimated ± 12.0 pp, low confidence
DeepSeek V3 03240.0%0.0%estimated ± 12.0 pp, low confidence
Gemini 2.5 Pro0.0%0.0%estimated ± 12.0 pp, low confidence
Gemini 3.1 Pro14.7%9.7%estimated ± 12.0 pp, low confidence
Gemini 3.5 Flash-Lite24.3%16.9%estimated ± 12.0 pp, low confidence
Gemini 3.6 Flash39.3%29.9%estimated ± 12.0 pp, low confidence
Gemma 3 27B0.0%0.0%estimated ± 12.0 pp, low confidence
Gemma 4 12B0.0%0.0%estimated ± 12.0 pp, low confidence
Gemma 4 26B A4B3.4%2.1%estimated ± 12.0 pp, low confidence
Gemma 4 31B6.1%3.8%estimated ± 12.0 pp, low confidence
Gemma 4 E2B0.0%0.0%estimated ± 12.0 pp, low confidence
Gemma 4 E4B0.0%0.0%estimated ± 12.0 pp, low confidence
GLM-4.725.8%18.1%estimated ± 12.0 pp, low confidence
GLM-5.131.0%22.4%estimated ± 12.0 pp, low confidence
GPT-4.1 mini0.0%0.0%estimated ± 12.0 pp, low confidence
GPT-4.1 nano0.0%0.0%estimated ± 12.0 pp, low confidence
GPT-4o0.0%0.0%estimated ± 12.0 pp, low confidence
GPT-4o mini0.0%0.0%estimated ± 12.0 pp, low confidence
GPT-5.116.5%11.0%estimated ± 12.0 pp, low confidence
GPT-5.437.4%28.1%estimated ± 12.0 pp, low confidence
GPT-5.4 mini25.8%18.1%estimated ± 12.0 pp, low confidence
GPT-5.4 nano22.5%15.5%estimated ± 12.0 pp, low confidence
GPT-5 (high)21.0%14.3%estimated ± 12.0 pp, low confidence
GPT-OSS 120B5.6%3.5%estimated ± 12.0 pp, low confidence
GPT-OSS 20B0.0%0.0%estimated ± 12.0 pp, low confidence
Granite 4.2 30B3.9%2.4%estimated ± 12.0 pp, low confidence
Granite 4.2 3B0.0%0.0%estimated ± 12.0 pp, low confidence
Granite 4.2 8B0.0%0.0%estimated ± 12.0 pp, low confidence
Grok 4.329.2%20.9%estimated ± 12.0 pp, low confidence
Grok 4.544.5%34.9%estimated ± 12.0 pp, low confidence
Hy328.3%20.1%estimated ± 12.0 pp, low confidence
Hy3 Preview35.8%26.6%estimated ± 12.0 pp, low confidence
K-Exaone0.0%0.0%estimated ± 12.0 pp, low confidence
Kimi K2.517.2%11.5%estimated ± 12.0 pp, low confidence
Kimi K2.5 (Reasoning)17.2%11.5%estimated ± 12.0 pp, low confidence
Kimi K2.7 Code27.0%19.1%estimated ± 12.0 pp, low confidence
LFM2.5-2.6B0.0%0.0%estimated ± 12.0 pp, low confidence
Ling 2.6 Flash0.0%0.0%estimated ± 12.0 pp, low confidence
Ling 3.0 Flash FP822.4%15.4%estimated ± 12.0 pp, low confidence
Ling 3.0 Flash VL33.2%24.3%estimated ± 12.0 pp, low confidence
Ling 3.0 Tiny3.2%2.0%estimated ± 12.0 pp, low confidence
Ling 3.1 Flash56.1%47.6%estimated ± 12.0 pp, low confidence
Llama 4 Maverick0.0%0.0%estimated ± 12.0 pp, low confidence
Llama 4 Scout0.0%0.0%estimated ± 12.0 pp, low confidence
Mercury 2.50.0%0.0%estimated ± 12.0 pp, low confidence
MiMo-V2.5-Pro31.2%22.6%estimated ± 12.0 pp, low confidence
MiMo-V2-Flash6.6%4.2%estimated ± 12.0 pp, low confidence
MiniCPM5-2B10.5%6.8%estimated ± 12.0 pp, low confidence
MiniMax M2.725.7%18.0%estimated ± 12.0 pp, low confidence
Mistral Large 30.0%0.0%estimated ± 12.0 pp, low confidence
Mistral Medium 3.5 128B13.2%8.6%estimated ± 12.0 pp, low confidence
Mistral Small 40.0%0.0%estimated ± 12.0 pp, low confidence
Mistral Small 4 (Reasoning)0.0%0.0%estimated ± 12.0 pp, low confidence
Muse Spark25.1%17.5%estimated ± 12.0 pp, low confidence
Muse Spark 1.248.9%39.5%estimated ± 12.0 pp, low confidence
Nemotron 3.5 Lightning 30B A3B NVFP47.0%4.4%estimated ± 12.0 pp, low confidence
Nemotron 3 Nano 30B0.0%0.0%estimated ± 12.0 pp, low confidence
Nemotron 3 Nano Omni 30B A3B0.0%0.0%estimated ± 12.0 pp, low confidence
Nemotron 3 Super 100B0.0%0.0%estimated ± 12.0 pp, low confidence
North Mini Code0.0%0.0%estimated ± 12.0 pp, low confidence
Quasar 438B34.2%25.2%estimated ± 12.0 pp, low confidence
Qwen3.5-122B-A10B15.8%10.5%estimated ± 12.0 pp, low confidence
Qwen3.6-27B24.4%17.0%estimated ± 12.0 pp, low confidence
Qwen3.6-35B-A3B19.8%13.4%estimated ± 12.0 pp, low confidence
Qwen3.6 Plus24.7%17.2%estimated ± 12.0 pp, low confidence
Qwen3.8 Max Preview58.6%50.6%estimated ± 12.0 pp, low confidence
Solar Pro 30.0%0.0%estimated ± 12.0 pp, low confidence
Solar Pro 430.5%22.0%estimated ± 12.0 pp, low confidence
Trinity-Large-Preview0.0%0.0%estimated ± 12.0 pp, low confidence
Trinity-Large-Thinking0.0%0.0%estimated ± 12.0 pp, low confidence
Ultravox v0.6 Llama 3.3 70B0.0%0.0%estimated ± 12.0 pp, low confidence