benchgap
Calibration

Gert Labs → CyberGym

CyberGym is estimated from Gert Labs with a inverse Michaelis–Menten curve fitted on 7 models measured on both: y = 1.79975·(x − 0.0974) / (0.0974 + 2.0000 − x), R² = 0.59, cross-validated error 10.0 pp. It is used for 18 estimates.

Estimated modelGert LabsCyberGymSource
Claude 4 Sonnet39.7%31.7%estimated ± 10.0 pp, low confidence
Claude Sonnet 4.548.5%43.3%estimated ± 10.0 pp, low confidence
DeepSeek V3.229.6%19.8%estimated ± 10.0 pp, low confidence
Gemini 3.1 Flash-Lite38.5%30.2%estimated ± 10.0 pp, low confidence
Gemini 3 Flash56.6%55.1%estimated ± 10.0 pp, low confidence
Gemini 3 Pro63.2%65.7%estimated ± 10.0 pp, low confidence
GLM-5V-Turbo30.8%21.1%estimated ± 10.0 pp, low confidence
GPT-4.125.7%15.6%estimated ± 10.0 pp, low confidence
GPT-5.1-Codex49.7%44.9%estimated ± 10.0 pp, low confidence
GPT-5.2-Codex51.8%47.9%estimated ± 10.0 pp, low confidence
GPT-5.3 Codex57.5%56.4%estimated ± 10.0 pp, low confidence
Grok 442.3%35.1%estimated ± 10.0 pp, low confidence
Grok 4.1 Fast47.3%41.6%estimated ± 10.0 pp, low confidence
Grok 4.2038.4%30.1%estimated ± 10.0 pp, low confidence
Grok Build 0.149.2%44.2%estimated ± 10.0 pp, low confidence
MiMo-V2.546.9%41.1%estimated ± 10.0 pp, low confidence
MiMo-V2-Pro36.7%28.0%estimated ± 10.0 pp, low confidence
Qwen3 Max43.7%36.9%estimated ± 10.0 pp, low confidence