benchgap
Calibration

GDPval-AA → JobBench

JobBench is estimated from GDPval-AA with a linear curve fitted on 13 models measured on both: y = 1.2398·x + -0.1001, R² = 0.83, cross-validated error 8.6 pp. It is used for 61 estimates.

Estimated modelGDPval-AAJobBenchSource
A.X K223.2%18.8%estimated ± 8.6 pp, medium confidence
Celeris-10.0%0.0%estimated ± 8.6 pp, low confidence
Claude Fable 555.6%58.9%estimated ± 8.6 pp, medium confidence
Claude Sonnet 548.3%49.9%estimated ± 8.6 pp, medium confidence
Command A+0.9%0.0%estimated ± 8.6 pp, low confidence
DeepSeek V30.0%0.0%estimated ± 8.6 pp, low confidence
DeepSeek V3 03240.0%0.0%estimated ± 8.6 pp, low confidence
DeepSeek V4.1 Flash55.0%58.2%estimated ± 8.6 pp, medium confidence
Gemini 2.5 Pro0.0%0.0%estimated ± 8.6 pp, low confidence
Gemini 3.5 Flash-Lite24.3%20.1%estimated ± 8.6 pp, medium confidence
Gemini 3.6 Flash39.3%38.7%estimated ± 8.6 pp, medium confidence
Gemma 3 27B0.0%0.0%estimated ± 8.6 pp, low confidence
Gemma 4 12B0.0%0.0%estimated ± 8.6 pp, low confidence
Gemma 4 26B A4B3.4%0.0%estimated ± 8.6 pp, low confidence
Gemma 4 31B6.1%0.0%estimated ± 8.6 pp, low confidence
Gemma 4 E2B0.0%0.0%estimated ± 8.6 pp, low confidence
Gemma 4 E4B0.0%0.0%estimated ± 8.6 pp, low confidence
GLM-4.725.8%22.0%estimated ± 8.6 pp, medium confidence
GPT-4.1 mini0.0%0.0%estimated ± 8.6 pp, low confidence
GPT-4.1 nano0.0%0.0%estimated ± 8.6 pp, low confidence
GPT-4o0.0%0.0%estimated ± 8.6 pp, low confidence
GPT-4o mini0.0%0.0%estimated ± 8.6 pp, low confidence
GPT-5.116.5%10.4%estimated ± 8.6 pp, low confidence
GPT-6.1 Sol53.8%56.7%estimated ± 8.6 pp, medium confidence
GPT-6 Luna46.9%48.1%estimated ± 8.6 pp, medium confidence
Granite 4.2 30B3.9%0.0%estimated ± 8.6 pp, low confidence
Granite 4.2 3B0.0%0.0%estimated ± 8.6 pp, low confidence
Granite 4.2 8B0.0%0.0%estimated ± 8.6 pp, low confidence
Grok 4.544.5%45.2%estimated ± 8.6 pp, medium confidence
Grok 4.656.1%59.5%estimated ± 8.6 pp, medium confidence
Grok 4.760.8%65.4%estimated ± 8.6 pp, low confidence
Hy328.3%25.1%estimated ± 8.6 pp, medium confidence
Hy3 Preview35.8%34.4%estimated ± 8.6 pp, medium confidence
K-Exaone0.0%0.0%estimated ± 8.6 pp, low confidence
Ling 2.6 Flash0.0%0.0%estimated ± 8.6 pp, low confidence
Ling 3.0 Flash FP822.4%17.8%estimated ± 8.6 pp, medium confidence
Ling 3.0 Flash VL33.2%31.1%estimated ± 8.6 pp, medium confidence
Ling 3.0 Tiny3.2%0.0%estimated ± 8.6 pp, low confidence
Ling 3.1 Flash56.1%59.5%estimated ± 8.6 pp, medium confidence
Llama 4 Maverick0.0%0.0%estimated ± 8.6 pp, low confidence
Llama 4 Scout0.0%0.0%estimated ± 8.6 pp, low confidence
Mercury 2.50.0%0.0%estimated ± 8.6 pp, low confidence
MiMo-V2-Flash6.6%0.0%estimated ± 8.6 pp, low confidence
MiniCPM5-2B10.5%3.0%estimated ± 8.6 pp, low confidence
Mistral Large 30.0%0.0%estimated ± 8.6 pp, low confidence
Mistral Large 446.2%47.3%estimated ± 8.6 pp, medium confidence
Mistral Medium 3.5 128B13.2%6.4%estimated ± 8.6 pp, low confidence
Mistral Small 40.0%0.0%estimated ± 8.6 pp, low confidence
Mistral Small 4 (Reasoning)0.0%0.0%estimated ± 8.6 pp, low confidence
Muse Spark 1.248.9%50.6%estimated ± 8.6 pp, medium confidence
Nemotron 3.5 Lightning 30B A3B NVFP47.0%0.0%estimated ± 8.6 pp, low confidence
Nemotron 3 Nano 30B0.0%0.0%estimated ± 8.6 pp, low confidence
Nemotron 3 Nano Omni 30B A3B0.0%0.0%estimated ± 8.6 pp, low confidence
Nemotron 3 Ultra33.1%31.0%estimated ± 8.6 pp, medium confidence
North Mini Code0.0%0.0%estimated ± 8.6 pp, low confidence
Quasar 438B34.2%32.4%estimated ± 8.6 pp, medium confidence
Qwen3.5-122B-A10B15.8%9.6%estimated ± 8.6 pp, low confidence
Solar Pro 30.0%0.0%estimated ± 8.6 pp, low confidence
Trinity-Large-Preview0.0%0.0%estimated ± 8.6 pp, low confidence
Trinity-Large-Thinking0.0%0.0%estimated ± 8.6 pp, low confidence
Ultravox v0.6 Llama 3.3 70B0.0%0.0%estimated ± 8.6 pp, low confidence