benchgap
Calibration

GDPval-AA → Gert Labs

Gert Labs is estimated from GDPval-AA with a Hill curve fitted on 25 models measured on both: y = 0.3597 + (1.2000 − 0.3597)·x^2.11 / (0.56360^2.11 + x^2.11), R² = 0.60, cross-validated error 8.9 pp. It is used for 31 estimates.

Estimated modelGDPval-AAGert LabsSource
A.X K223.2%47.2%estimated ± 8.9 pp, medium confidence
Apodex 1.134.8%58.3%estimated ± 8.9 pp, medium confidence
Apodex 1.1 Mini34.8%58.3%estimated ± 8.9 pp, medium confidence
Claude Sonnet 5.567.0%85.6%estimated ± 8.9 pp, low confidence
DeepSeek V3 03240.0%36.0%estimated ± 8.9 pp, medium confidence
Gemma 4 12B0.0%36.0%estimated ± 8.9 pp, medium confidence
Gemma 4 26B A4B3.4%36.2%estimated ± 8.9 pp, medium confidence
Gemma 4 E2B0.0%36.0%estimated ± 8.9 pp, medium confidence
Gemma 4 E4B0.0%36.0%estimated ± 8.9 pp, medium confidence
GPT-4.1 mini0.0%36.0%estimated ± 8.9 pp, medium confidence
GPT-4.1 nano0.0%36.0%estimated ± 8.9 pp, medium confidence
GPT-4o0.0%36.0%estimated ± 8.9 pp, medium confidence
GPT-4o mini0.0%36.0%estimated ± 8.9 pp, medium confidence
GPT-6.1 Sol53.8%75.9%estimated ± 8.9 pp, low confidence
GPT-6 Luna46.9%70.0%estimated ± 8.9 pp, medium confidence
Granite 4.2 30B3.9%36.3%estimated ± 8.9 pp, medium confidence
Granite 4.2 3B0.0%36.0%estimated ± 8.9 pp, medium confidence
Grok 4.760.8%81.3%estimated ± 8.9 pp, low confidence
K-Exaone0.0%36.0%estimated ± 8.9 pp, medium confidence
LFM2.5-2.6B0.0%36.0%estimated ± 8.9 pp, medium confidence
Ling 2.6 Flash0.0%36.0%estimated ± 8.9 pp, medium confidence
Ling 3.0 Flash VL33.2%56.7%estimated ± 8.9 pp, medium confidence
Ling 3.0 Tiny3.2%36.2%estimated ± 8.9 pp, medium confidence
Mercury 2.50.0%36.0%estimated ± 8.9 pp, medium confidence
MiMo-V2-Flash6.6%36.9%estimated ± 8.9 pp, medium confidence
MiniCPM5-2B10.5%38.3%estimated ± 8.9 pp, medium confidence
Mistral Large 446.2%69.3%estimated ± 8.9 pp, medium confidence
Nemotron 3 Nano Omni 30B A3B0.0%36.0%estimated ± 8.9 pp, medium confidence
North Mini Code0.0%36.0%estimated ± 8.9 pp, medium confidence
Solar Pro 30.0%36.0%estimated ± 8.9 pp, medium confidence
Ultravox v0.6 Llama 3.3 70B0.0%36.0%estimated ± 8.9 pp, medium confidence