Calibration
Gert Labs → CyberGym
CyberGym is estimated from Gert Labs with a inverse Michaelis–Menten curve fitted on 7 models measured on both: y = 1.79975·(x − 0.0974) / (0.0974 + 2.0000 − x), R² = 0.59, cross-validated error 10.0 pp. It is used for 18 estimates.
| Estimated model | Gert Labs | CyberGym | Source |
|---|---|---|---|
| Claude 4 Sonnet | 39.7% | 31.7% | estimated ± 10.0 pp, low confidence |
| Claude Sonnet 4.5 | 48.5% | 43.3% | estimated ± 10.0 pp, low confidence |
| DeepSeek V3.2 | 29.6% | 19.8% | estimated ± 10.0 pp, low confidence |
| Gemini 3.1 Flash-Lite | 38.5% | 30.2% | estimated ± 10.0 pp, low confidence |
| Gemini 3 Flash | 56.6% | 55.1% | estimated ± 10.0 pp, low confidence |
| Gemini 3 Pro | 63.2% | 65.7% | estimated ± 10.0 pp, low confidence |
| GLM-5V-Turbo | 30.8% | 21.1% | estimated ± 10.0 pp, low confidence |
| GPT-4.1 | 25.7% | 15.6% | estimated ± 10.0 pp, low confidence |
| GPT-5.1-Codex | 49.7% | 44.9% | estimated ± 10.0 pp, low confidence |
| GPT-5.2-Codex | 51.8% | 47.9% | estimated ± 10.0 pp, low confidence |
| GPT-5.3 Codex | 57.5% | 56.4% | estimated ± 10.0 pp, low confidence |
| Grok 4 | 42.3% | 35.1% | estimated ± 10.0 pp, low confidence |
| Grok 4.1 Fast | 47.3% | 41.6% | estimated ± 10.0 pp, low confidence |
| Grok 4.20 | 38.4% | 30.1% | estimated ± 10.0 pp, low confidence |
| Grok Build 0.1 | 49.2% | 44.2% | estimated ± 10.0 pp, low confidence |
| MiMo-V2.5 | 46.9% | 41.1% | estimated ± 10.0 pp, low confidence |
| MiMo-V2-Pro | 36.7% | 28.0% | estimated ± 10.0 pp, low confidence |
| Qwen3 Max | 43.7% | 36.9% | estimated ± 10.0 pp, low confidence |