benchgap
Calibration

GDPval-AA → DRACO

DRACO is estimated from GDPval-AA with a linear curve fitted on 5 models measured on both: y = 0.4102·x + 0.6154, R² = 0.96, cross-validated error 3.0 pp. It is used for 117 estimates.

Estimated modelGDPval-AADRACOSource
A.X K223.2%71.1%estimated ± 3.0 pp, medium confidence
Apodex 1.134.8%75.8%estimated ± 3.0 pp, medium confidence
Apodex 1.1 Mini34.8%75.8%estimated ± 3.0 pp, medium confidence
Celeris-10.0%61.5%estimated ± 3.0 pp, low confidence
Claude Fable 555.6%84.4%estimated ± 3.0 pp, medium confidence
Claude Fable 5.162.9%87.3%estimated ± 3.0 pp, medium confidence
Claude Opus 4.7 (Adaptive)42.8%79.1%estimated ± 3.0 pp, medium confidence
Claude Opus 4.847.8%81.2%estimated ± 3.0 pp, medium confidence
Claude Opus 5.568.3%89.6%estimated ± 3.0 pp, low confidence
Claude Sonnet 548.3%81.4%estimated ± 3.0 pp, medium confidence
Command A+0.9%61.9%estimated ± 3.0 pp, low confidence
DeepSeek V30.0%61.5%estimated ± 3.0 pp, low confidence
DeepSeek V3 03240.0%61.5%estimated ± 3.0 pp, low confidence
DeepSeek V4.1 Flash55.0%84.1%estimated ± 3.0 pp, medium confidence
DeepSeek V4 Flash 073146.9%80.8%estimated ± 3.0 pp, medium confidence
DeepSeek V4 Pro 081354.5%83.9%estimated ± 3.0 pp, medium confidence
Gemini 2.5 Pro0.0%61.5%estimated ± 3.0 pp, low confidence
Gemini 3.1 Pro14.7%67.6%estimated ± 3.0 pp, low confidence
Gemini 3.5 Flash42.2%78.9%estimated ± 3.0 pp, medium confidence
Gemini 3.5 Flash-Lite24.3%71.5%estimated ± 3.0 pp, medium confidence
Gemini 3.6 Flash39.3%77.7%estimated ± 3.0 pp, medium confidence
Gemini 3.7 Flash44.6%79.8%estimated ± 3.0 pp, medium confidence
Gemini 3.8 Flash46.8%80.7%estimated ± 3.0 pp, medium confidence
Gemini 4 Argon56.3%84.6%estimated ± 3.0 pp, medium confidence
Gemma 3 27B0.0%61.5%estimated ± 3.0 pp, low confidence
Gemma 4 12B0.0%61.5%estimated ± 3.0 pp, low confidence
Gemma 4 26B A4B3.4%62.9%estimated ± 3.0 pp, low confidence
Gemma 4 31B6.1%64.0%estimated ± 3.0 pp, low confidence
Gemma 4 E2B0.0%61.5%estimated ± 3.0 pp, low confidence
Gemma 4 E4B0.0%61.5%estimated ± 3.0 pp, low confidence
GLM-4.725.8%72.1%estimated ± 3.0 pp, medium confidence
GLM-5.131.0%74.3%estimated ± 3.0 pp, medium confidence
GLM-5.243.7%79.5%estimated ± 3.0 pp, medium confidence
GLM-5.357.6%85.2%estimated ± 3.0 pp, medium confidence
GPT-4.1 mini0.0%61.5%estimated ± 3.0 pp, low confidence
GPT-4.1 nano0.0%61.5%estimated ± 3.0 pp, low confidence
GPT-4o0.0%61.5%estimated ± 3.0 pp, low confidence
GPT-4o mini0.0%61.5%estimated ± 3.0 pp, low confidence
GPT-5.116.5%68.3%estimated ± 3.0 pp, low confidence
GPT-5.437.4%76.9%estimated ± 3.0 pp, medium confidence
GPT-5.4 mini25.8%72.1%estimated ± 3.0 pp, medium confidence
GPT-5.4 nano22.5%70.8%estimated ± 3.0 pp, medium confidence
GPT-5.542.7%79.1%estimated ± 3.0 pp, medium confidence
GPT-5.6 Luna48.2%81.3%estimated ± 3.0 pp, medium confidence
GPT-5.6 Sol55.6%84.4%estimated ± 3.0 pp, medium confidence
GPT-5.6 Terra47.7%81.1%estimated ± 3.0 pp, medium confidence
GPT-5 (high)21.0%70.2%estimated ± 3.0 pp, low confidence
GPT-6.1 Sol53.8%83.6%estimated ± 3.0 pp, medium confidence
GPT-6 Astra52.1%82.9%estimated ± 3.0 pp, medium confidence
GPT-6 Luna46.9%80.8%estimated ± 3.0 pp, medium confidence
GPT-6 Sol50.5%82.3%estimated ± 3.0 pp, medium confidence
GPT-OSS 120B5.6%63.8%estimated ± 3.0 pp, low confidence
GPT-OSS 20B0.0%61.5%estimated ± 3.0 pp, low confidence
Granite 4.2 30B3.9%63.1%estimated ± 3.0 pp, low confidence
Granite 4.2 3B0.0%61.5%estimated ± 3.0 pp, low confidence
Granite 4.2 8B0.0%61.5%estimated ± 3.0 pp, low confidence
Grok 4.329.2%73.5%estimated ± 3.0 pp, medium confidence
Grok 4.544.5%79.8%estimated ± 3.0 pp, medium confidence
Grok 4.656.1%84.6%estimated ± 3.0 pp, medium confidence
Grok 4.760.8%86.5%estimated ± 3.0 pp, medium confidence
Hy328.3%73.2%estimated ± 3.0 pp, medium confidence
Hy3 Preview35.8%76.2%estimated ± 3.0 pp, medium confidence
Inkling28.9%73.4%estimated ± 3.0 pp, medium confidence
Inkling-Small31.2%74.3%estimated ± 3.0 pp, medium confidence
K-Exaone0.0%61.5%estimated ± 3.0 pp, low confidence
Kimi K2.627.0%72.6%estimated ± 3.0 pp, medium confidence
Kimi K2.517.2%68.6%estimated ± 3.0 pp, low confidence
Kimi K2.5 (Reasoning)17.2%68.6%estimated ± 3.0 pp, low confidence
Kimi K2.7 Code27.0%72.6%estimated ± 3.0 pp, medium confidence
Kimi K351.8%82.8%estimated ± 3.0 pp, medium confidence
LFM2.5-2.6B0.0%61.5%estimated ± 3.0 pp, low confidence
Ling 2.6 Flash0.0%61.5%estimated ± 3.0 pp, low confidence
Ling 3.0 Flash FP822.4%70.7%estimated ± 3.0 pp, medium confidence
Ling 3.0 Flash VL33.2%75.2%estimated ± 3.0 pp, medium confidence
Ling 3.0 Tiny3.2%62.9%estimated ± 3.0 pp, low confidence
Llama 4 Maverick0.0%61.5%estimated ± 3.0 pp, low confidence
Llama 4 Scout0.0%61.5%estimated ± 3.0 pp, low confidence
Mercury 2.50.0%61.5%estimated ± 3.0 pp, low confidence
MiMo-V2.5-Pro31.2%74.3%estimated ± 3.0 pp, medium confidence
MiMo-V2.6-Flash55.5%84.3%estimated ± 3.0 pp, medium confidence
MiMo-V2.6-Pro59.3%85.9%estimated ± 3.0 pp, medium confidence
MiMo-V2-Flash6.6%64.3%estimated ± 3.0 pp, low confidence
MiniCPM5-2B10.5%65.9%estimated ± 3.0 pp, low confidence
MiniMax M2.725.7%72.1%estimated ± 3.0 pp, medium confidence
MiniMax M337.3%76.8%estimated ± 3.0 pp, medium confidence
Mistral Large 30.0%61.5%estimated ± 3.0 pp, low confidence
Mistral Large 446.2%80.5%estimated ± 3.0 pp, medium confidence
Mistral Medium 3.5 128B13.2%67.0%estimated ± 3.0 pp, low confidence
Mistral Small 40.0%61.5%estimated ± 3.0 pp, low confidence
Mistral Small 4 (Reasoning)0.0%61.5%estimated ± 3.0 pp, low confidence
Muse Glimmer 30B14.5%67.5%estimated ± 3.0 pp, low confidence
Muse Spark25.1%71.8%estimated ± 3.0 pp, medium confidence
Muse Spark 1.135.7%76.2%estimated ± 3.0 pp, medium confidence
Muse Spark 1.248.9%81.6%estimated ± 3.0 pp, medium confidence
Muse Spark 1.359.2%85.8%estimated ± 3.0 pp, medium confidence
Nemotron 3.5 Lightning 30B A3B NVFP47.0%64.4%estimated ± 3.0 pp, low confidence
Nemotron 3 Nano 30B0.0%61.5%estimated ± 3.0 pp, low confidence
Nemotron 3 Nano Omni 30B A3B0.0%61.5%estimated ± 3.0 pp, low confidence
Nemotron 3 Super 100B0.0%61.5%estimated ± 3.0 pp, low confidence
Nemotron 3 Ultra33.1%75.1%estimated ± 3.0 pp, medium confidence
North Mini Code0.0%61.5%estimated ± 3.0 pp, low confidence
Quasar 438B34.2%75.6%estimated ± 3.0 pp, medium confidence
Qwen3.5-122B-A10B15.8%68.0%estimated ± 3.0 pp, low confidence
Qwen3.6-27B24.4%71.6%estimated ± 3.0 pp, medium confidence
Qwen3.6-35B-A3B19.8%69.7%estimated ± 3.0 pp, low confidence
Qwen3.6 Plus24.7%71.7%estimated ± 3.0 pp, medium confidence
Qwen3.7 Max31.6%74.5%estimated ± 3.0 pp, medium confidence
Qwen3.7 Plus13.5%67.1%estimated ± 3.0 pp, low confidence
Qwen3.8-27B46.2%80.5%estimated ± 3.0 pp, medium confidence
Qwen3.8-Flash-Next56.6%84.8%estimated ± 3.0 pp, medium confidence
Qwen3.8 Max Preview58.6%85.6%estimated ± 3.0 pp, medium confidence
Solar Pro 30.0%61.5%estimated ± 3.0 pp, low confidence
Solar Pro 430.5%74.1%estimated ± 3.0 pp, medium confidence
Step 3.7 Flash25.8%72.1%estimated ± 3.0 pp, medium confidence
Trinity-Large-Preview0.0%61.5%estimated ± 3.0 pp, low confidence
Trinity-Large-Thinking0.0%61.5%estimated ± 3.0 pp, low confidence
Ultravox v0.6 Llama 3.3 70B0.0%61.5%estimated ± 3.0 pp, low confidence