benchgap
Calibration

Gert Labs → Toolathlon

Toolathlon is estimated from Gert Labs with a Michaelis–Menten + offset curve fitted on 13 models measured on both: y = 0.0000 + 2.0000·x / (1.89254 + x), R² = 0.58, cross-validated error 7.2 pp. It is used for 16 estimates.

Estimated modelGert LabsToolathlonSource
Claude 4 Sonnet39.7%34.7%estimated ± 7.2 pp, low confidence
Claude Sonnet 4.548.5%40.8%estimated ± 7.2 pp, medium confidence
DeepSeek V3.229.6%27.0%estimated ± 7.2 pp, low confidence
Gemini 3.1 Flash-Lite38.5%33.8%estimated ± 7.2 pp, low confidence
Gemini 3 Flash56.6%46.1%estimated ± 7.2 pp, medium confidence
Gemini 3 Pro63.2%50.1%estimated ± 7.2 pp, medium confidence
GLM-5V-Turbo30.8%28.0%estimated ± 7.2 pp, low confidence
GPT-4.125.7%23.9%estimated ± 7.2 pp, low confidence
GPT-5.1-Codex49.7%41.6%estimated ± 7.2 pp, medium confidence
GPT-5.2-Codex51.8%43.0%estimated ± 7.2 pp, medium confidence
GPT-5.3 Codex57.5%46.6%estimated ± 7.2 pp, medium confidence
Grok 442.3%36.6%estimated ± 7.2 pp, medium confidence
Grok 4.1 Fast47.3%40.0%estimated ± 7.2 pp, medium confidence
Grok 4.2038.4%33.7%estimated ± 7.2 pp, low confidence
Grok Build 0.149.2%41.2%estimated ± 7.2 pp, medium confidence
Qwen3 Max43.7%37.5%estimated ± 7.2 pp, medium confidence