benchgap
Calibration

GDPval-AA → DeepSearchQA

DeepSearchQA is estimated from GDPval-AA with a offset logistic curve fitted on 14 models measured on both: y = 0.0000 + (0.9055 − 0.0000) / (1 + exp(−12.62·(x − 0.0386))), R² = 0.85, cross-validated error 8.4 pp. It is used for 40 estimates.

Estimated modelGDPval-AADeepSearchQASource
A.X K223.2%83.3%estimated ± 8.4 pp, medium confidence
Apodex 1.1 Mini34.8%88.8%estimated ± 8.4 pp, medium confidence
Claude Opus 5.568.3%90.5%estimated ± 8.4 pp, low confidence
Claude Sonnet 5.567.0%90.5%estimated ± 8.4 pp, low confidence
DeepSeek V3 03240.0%34.5%estimated ± 8.4 pp, medium confidence
DeepSeek V4.1 Flash55.0%90.4%estimated ± 8.4 pp, medium confidence
Gemini 4 Argon56.3%90.4%estimated ± 8.4 pp, medium confidence
Gemma 4 12B0.0%34.5%estimated ± 8.4 pp, medium confidence
Gemma 4 26B A4B3.4%44.0%estimated ± 8.4 pp, medium confidence
Gemma 4 E2B0.0%34.5%estimated ± 8.4 pp, medium confidence
Gemma 4 E4B0.0%34.5%estimated ± 8.4 pp, medium confidence
GPT-4.1 mini0.0%34.5%estimated ± 8.4 pp, medium confidence
GPT-4.1 nano0.0%34.5%estimated ± 8.4 pp, medium confidence
GPT-4o0.0%34.5%estimated ± 8.4 pp, medium confidence
GPT-4o mini0.0%34.5%estimated ± 8.4 pp, medium confidence
GPT-5.116.5%75.3%estimated ± 8.4 pp, medium confidence
GPT-5 (high)21.0%81.2%estimated ± 8.4 pp, medium confidence
GPT-6.1 Sol53.8%90.4%estimated ± 8.4 pp, medium confidence
GPT-6 Luna46.9%90.2%estimated ± 8.4 pp, medium confidence
GPT-6 Sol50.5%90.3%estimated ± 8.4 pp, medium confidence
Granite 4.2 30B3.9%45.4%estimated ± 8.4 pp, medium confidence
Granite 4.2 3B0.0%34.5%estimated ± 8.4 pp, medium confidence
Grok 4.760.8%90.5%estimated ± 8.4 pp, medium confidence
K-Exaone0.0%34.5%estimated ± 8.4 pp, medium confidence
LFM2.5-2.6B0.0%34.5%estimated ± 8.4 pp, medium confidence
Ling 2.6 Flash0.0%34.5%estimated ± 8.4 pp, medium confidence
Ling 3.0 Flash VL33.2%88.4%estimated ± 8.4 pp, medium confidence
Ling 3.0 Tiny3.2%43.4%estimated ± 8.4 pp, medium confidence
Ling 3.1 Flash56.1%90.4%estimated ± 8.4 pp, medium confidence
MiMo-V2.6-Flash55.5%90.4%estimated ± 8.4 pp, medium confidence
MiMo-V2.6-Pro59.3%90.5%estimated ± 8.4 pp, medium confidence
MiMo-V2-Flash6.6%53.0%estimated ± 8.4 pp, medium confidence
MiniCPM5-2B10.5%63.2%estimated ± 8.4 pp, medium confidence
Mistral Large 446.2%90.1%estimated ± 8.4 pp, medium confidence
Nemotron 3 Nano Omni 30B A3B0.0%34.5%estimated ± 8.4 pp, medium confidence
North Mini Code0.0%34.5%estimated ± 8.4 pp, medium confidence
Qwen3.6 Plus24.7%84.5%estimated ± 8.4 pp, medium confidence
Qwen3.8-Flash-Next56.6%90.4%estimated ± 8.4 pp, medium confidence
Solar Pro 30.0%34.5%estimated ± 8.4 pp, medium confidence
Ultravox v0.6 Llama 3.3 70B0.0%34.5%estimated ± 8.4 pp, medium confidence