benchgap
Calibration

GDPval-AA → Toolathlon

Toolathlon is estimated from GDPval-AA with a offset logistic curve fitted on 19 models measured on both: y = 0.2691 + (0.5585 − 0.2691) / (1 + exp(−58.59·(x − 0.2463))), R² = 0.74, cross-validated error 6.3 pp. It is used for 73 estimates.

Estimated modelGDPval-AAToolathlonSource
A.X K223.2%35.6%estimated ± 6.3 pp, medium confidence
Apodex 1.134.8%55.8%estimated ± 6.3 pp, medium confidence
Apodex 1.1 Mini34.8%55.8%estimated ± 6.3 pp, medium confidence
Celeris-10.0%26.9%estimated ± 6.3 pp, low confidence
Claude Opus 5.568.3%55.9%estimated ± 6.3 pp, low confidence
Claude Sonnet 5.567.0%55.9%estimated ± 6.3 pp, low confidence
Command A+0.9%26.9%estimated ± 6.3 pp, low confidence
DeepSeek V30.0%26.9%estimated ± 6.3 pp, low confidence
DeepSeek V3 03240.0%26.9%estimated ± 6.3 pp, low confidence
DeepSeek V4.1 Flash55.0%55.9%estimated ± 6.3 pp, medium confidence
Gemini 2.5 Pro0.0%26.9%estimated ± 6.3 pp, low confidence
Gemini 3.1 Pro14.7%27.0%estimated ± 6.3 pp, low confidence
Gemini 3.5 Flash-Lite24.3%40.0%estimated ± 6.3 pp, medium confidence
Gemini 3.6 Flash39.3%55.8%estimated ± 6.3 pp, medium confidence
Gemini 4 Argon56.3%55.9%estimated ± 6.3 pp, low confidence
Gemma 3 27B0.0%26.9%estimated ± 6.3 pp, low confidence
Gemma 4 12B0.0%26.9%estimated ± 6.3 pp, low confidence
Gemma 4 26B A4B3.4%26.9%estimated ± 6.3 pp, low confidence
Gemma 4 31B6.1%26.9%estimated ± 6.3 pp, low confidence
Gemma 4 E2B0.0%26.9%estimated ± 6.3 pp, low confidence
Gemma 4 E4B0.0%26.9%estimated ± 6.3 pp, low confidence
GLM-5.357.6%55.9%estimated ± 6.3 pp, low confidence
GPT-4.1 mini0.0%26.9%estimated ± 6.3 pp, low confidence
GPT-4.1 nano0.0%26.9%estimated ± 6.3 pp, low confidence
GPT-4o0.0%26.9%estimated ± 6.3 pp, low confidence
GPT-4o mini0.0%26.9%estimated ± 6.3 pp, low confidence
GPT-5.116.5%27.2%estimated ± 6.3 pp, low confidence
GPT-5 (high)21.0%30.0%estimated ± 6.3 pp, medium confidence
GPT-6.1 Sol53.8%55.9%estimated ± 6.3 pp, medium confidence
GPT-6 Luna46.9%55.9%estimated ± 6.3 pp, medium confidence
GPT-6 Sol50.5%55.9%estimated ± 6.3 pp, medium confidence
GPT-OSS 120B5.6%26.9%estimated ± 6.3 pp, low confidence
GPT-OSS 20B0.0%26.9%estimated ± 6.3 pp, low confidence
Grok 4.329.2%54.0%estimated ± 6.3 pp, medium confidence
Grok 4.544.5%55.9%estimated ± 6.3 pp, medium confidence
Grok 4.760.8%55.9%estimated ± 6.3 pp, low confidence
Hy328.3%52.8%estimated ± 6.3 pp, medium confidence
Hy3 Preview35.8%55.8%estimated ± 6.3 pp, medium confidence
K-Exaone0.0%26.9%estimated ± 6.3 pp, low confidence
Kimi K2.7 Code27.0%50.1%estimated ± 6.3 pp, medium confidence
Ling 2.6 Flash0.0%26.9%estimated ± 6.3 pp, low confidence
Ling 3.0 Flash FP822.4%33.1%estimated ± 6.3 pp, medium confidence
Ling 3.0 Flash VL33.2%55.7%estimated ± 6.3 pp, medium confidence
Ling 3.0 Tiny3.2%26.9%estimated ± 6.3 pp, low confidence
Ling 3.1 Flash56.1%55.9%estimated ± 6.3 pp, low confidence
Llama 4 Maverick0.0%26.9%estimated ± 6.3 pp, low confidence
Llama 4 Scout0.0%26.9%estimated ± 6.3 pp, low confidence
MiMo-V2.6-Flash55.5%55.9%estimated ± 6.3 pp, medium confidence
MiMo-V2.6-Pro59.3%55.9%estimated ± 6.3 pp, low confidence
MiMo-V2-Flash6.6%26.9%estimated ± 6.3 pp, low confidence
MiniCPM5-2B10.5%26.9%estimated ± 6.3 pp, low confidence
Mistral Large 30.0%26.9%estimated ± 6.3 pp, low confidence
Mistral Large 446.2%55.9%estimated ± 6.3 pp, medium confidence
Mistral Small 40.0%26.9%estimated ± 6.3 pp, low confidence
Mistral Small 4 (Reasoning)0.0%26.9%estimated ± 6.3 pp, low confidence
Muse Glimmer 30B14.5%27.0%estimated ± 6.3 pp, low confidence
Muse Spark25.1%43.3%estimated ± 6.3 pp, medium confidence
Muse Spark 1.248.9%55.9%estimated ± 6.3 pp, medium confidence
Nemotron 3 Nano 30B0.0%26.9%estimated ± 6.3 pp, low confidence
Nemotron 3 Nano Omni 30B A3B0.0%26.9%estimated ± 6.3 pp, low confidence
Nemotron 3 Super 100B0.0%26.9%estimated ± 6.3 pp, low confidence
North Mini Code0.0%26.9%estimated ± 6.3 pp, low confidence
Quasar 438B34.2%55.7%estimated ± 6.3 pp, medium confidence
Qwen3.6-27B24.4%40.4%estimated ± 6.3 pp, medium confidence
Qwen3.7 Max31.6%55.4%estimated ± 6.3 pp, medium confidence
Qwen3.7 Plus13.5%27.0%estimated ± 6.3 pp, low confidence
Qwen3.8-27B46.2%55.9%estimated ± 6.3 pp, medium confidence
Qwen3.8-Flash-Next56.6%55.9%estimated ± 6.3 pp, low confidence
Qwen3.8 Max Preview58.6%55.9%estimated ± 6.3 pp, low confidence
Solar Pro 30.0%26.9%estimated ± 6.3 pp, low confidence
Trinity-Large-Preview0.0%26.9%estimated ± 6.3 pp, low confidence
Trinity-Large-Thinking0.0%26.9%estimated ± 6.3 pp, low confidence
Ultravox v0.6 Llama 3.3 70B0.0%26.9%estimated ± 6.3 pp, low confidence