benchgap
Calibration

GDPval-AA → OSWorld-Verified

OSWorld-Verified is estimated from GDPval-AA with a linear curve fitted on 20 models measured on both: y = 0.5249·x + 0.5555, R² = 0.46, cross-validated error 8.7 pp. It is used for 51 estimates.

Estimated modelGDPval-AAOSWorld-VerifiedSource
A.X K223.2%67.7%estimated ± 8.7 pp, low confidence
Apodex 1.1 Mini34.8%73.8%estimated ± 8.7 pp, low confidence
Celeris-10.0%55.6%estimated ± 8.7 pp, low confidence
Claude Sonnet 5.567.0%90.7%estimated ± 8.7 pp, low confidence
Command A+0.9%56.0%estimated ± 8.7 pp, low confidence
DeepSeek V30.0%55.6%estimated ± 8.7 pp, low confidence
DeepSeek V3 03240.0%55.6%estimated ± 8.7 pp, low confidence
Gemma 3 27B0.0%55.6%estimated ± 8.7 pp, low confidence
Gemma 4 12B0.0%55.6%estimated ± 8.7 pp, low confidence
Gemma 4 26B A4B3.4%57.3%estimated ± 8.7 pp, low confidence
Gemma 4 E2B0.0%55.6%estimated ± 8.7 pp, low confidence
Gemma 4 E4B0.0%55.6%estimated ± 8.7 pp, low confidence
GLM-5.243.7%78.5%estimated ± 8.7 pp, low confidence
GPT-4.1 mini0.0%55.6%estimated ± 8.7 pp, low confidence
GPT-4.1 nano0.0%55.6%estimated ± 8.7 pp, low confidence
GPT-4o0.0%55.6%estimated ± 8.7 pp, low confidence
GPT-4o mini0.0%55.6%estimated ± 8.7 pp, low confidence
GPT-6.1 Sol53.8%83.8%estimated ± 8.7 pp, low confidence
GPT-6 Luna46.9%80.2%estimated ± 8.7 pp, low confidence
GPT-OSS 20B0.0%55.6%estimated ± 8.7 pp, low confidence
Granite 4.2 30B3.9%57.6%estimated ± 8.7 pp, low confidence
Granite 4.2 3B0.0%55.6%estimated ± 8.7 pp, low confidence
Granite 4.2 8B0.0%55.6%estimated ± 8.7 pp, low confidence
Grok 4.544.5%78.9%estimated ± 8.7 pp, low confidence
Grok 4.760.8%87.5%estimated ± 8.7 pp, low confidence
Hy328.3%70.4%estimated ± 8.7 pp, low confidence
K-Exaone0.0%55.6%estimated ± 8.7 pp, low confidence
Kimi K2.7 Code27.0%69.7%estimated ± 8.7 pp, low confidence
LFM2.5-2.6B0.0%55.6%estimated ± 8.7 pp, low confidence
Ling 2.6 Flash0.0%55.6%estimated ± 8.7 pp, low confidence
Ling 3.0 Flash FP822.4%67.3%estimated ± 8.7 pp, low confidence
Ling 3.0 Flash VL33.2%73.0%estimated ± 8.7 pp, low confidence
Ling 3.0 Tiny3.2%57.2%estimated ± 8.7 pp, low confidence
Llama 4 Maverick0.0%55.6%estimated ± 8.7 pp, low confidence
Llama 4 Scout0.0%55.6%estimated ± 8.7 pp, low confidence
MiMo-V2-Flash6.6%59.0%estimated ± 8.7 pp, low confidence
MiniCPM5-2B10.5%61.1%estimated ± 8.7 pp, low confidence
Mistral Large 30.0%55.6%estimated ± 8.7 pp, low confidence
Mistral Large 446.2%79.8%estimated ± 8.7 pp, low confidence
Mistral Small 40.0%55.6%estimated ± 8.7 pp, low confidence
Mistral Small 4 (Reasoning)0.0%55.6%estimated ± 8.7 pp, low confidence
Muse Spark 1.248.9%81.2%estimated ± 8.7 pp, low confidence
Nemotron 3 Nano 30B0.0%55.6%estimated ± 8.7 pp, low confidence
Nemotron 3 Nano Omni 30B A3B0.0%55.6%estimated ± 8.7 pp, low confidence
Nemotron 3 Super 100B0.0%55.6%estimated ± 8.7 pp, low confidence
North Mini Code0.0%55.6%estimated ± 8.7 pp, low confidence
Quasar 438B34.2%73.5%estimated ± 8.7 pp, low confidence
Qwen3.8 Max Preview58.6%86.3%estimated ± 8.7 pp, low confidence
Solar Pro 30.0%55.6%estimated ± 8.7 pp, low confidence
Trinity-Large-Preview0.0%55.6%estimated ± 8.7 pp, low confidence
Ultravox v0.6 Llama 3.3 70B0.0%55.6%estimated ± 8.7 pp, low confidence