Calibration
Gert Labs → Toolathlon
Toolathlon is estimated from Gert Labs with a Michaelis–Menten + offset curve fitted on 13 models measured on both: y = 0.0000 + 2.0000·x / (1.89254 + x), R² = 0.58, cross-validated error 7.2 pp. It is used for 16 estimates.
| Estimated model | Gert Labs | Toolathlon | Source |
|---|---|---|---|
| Claude 4 Sonnet | 39.7% | 34.7% | estimated ± 7.2 pp, low confidence |
| Claude Sonnet 4.5 | 48.5% | 40.8% | estimated ± 7.2 pp, medium confidence |
| DeepSeek V3.2 | 29.6% | 27.0% | estimated ± 7.2 pp, low confidence |
| Gemini 3.1 Flash-Lite | 38.5% | 33.8% | estimated ± 7.2 pp, low confidence |
| Gemini 3 Flash | 56.6% | 46.1% | estimated ± 7.2 pp, medium confidence |
| Gemini 3 Pro | 63.2% | 50.1% | estimated ± 7.2 pp, medium confidence |
| GLM-5V-Turbo | 30.8% | 28.0% | estimated ± 7.2 pp, low confidence |
| GPT-4.1 | 25.7% | 23.9% | estimated ± 7.2 pp, low confidence |
| GPT-5.1-Codex | 49.7% | 41.6% | estimated ± 7.2 pp, medium confidence |
| GPT-5.2-Codex | 51.8% | 43.0% | estimated ± 7.2 pp, medium confidence |
| GPT-5.3 Codex | 57.5% | 46.6% | estimated ± 7.2 pp, medium confidence |
| Grok 4 | 42.3% | 36.6% | estimated ± 7.2 pp, medium confidence |
| Grok 4.1 Fast | 47.3% | 40.0% | estimated ± 7.2 pp, medium confidence |
| Grok 4.20 | 38.4% | 33.7% | estimated ± 7.2 pp, low confidence |
| Grok Build 0.1 | 49.2% | 41.2% | estimated ± 7.2 pp, medium confidence |
| Qwen3 Max | 43.7% | 37.5% | estimated ± 7.2 pp, medium confidence |