benchgap
Calibration

Vibe Code Bench → PostTrainBench v1.1

PostTrainBench v1.1 is estimated from Vibe Code Bench with a Hill curve fitted on 5 models measured on both: y = 0.1985 + (1.2000 − 0.1985)·x^6.00 / (1.10324^6.00 + x^6.00), R² = 0.89, cross-validated error 10.5 pp. It is used for 16 estimates.

Estimated modelVibe Code BenchPostTrainBench v1.1Source
Claude Haiku 4.5 Thinking11.4%19.9%estimated ± 10.5 pp, low confidence
Claude Opus 4.5 Thinking20.6%19.9%estimated ± 10.5 pp, low confidence
Claude Opus 4.6 (Adaptive)53.5%21.1%estimated ± 10.5 pp, low confidence
Claude Sonnet 4.5 Thinking22.6%19.9%estimated ± 10.5 pp, low confidence
DeepSeek V3.2 (Thinking)5.1%19.9%estimated ± 10.5 pp, low confidence
Gemini 3 Pro14.3%19.9%estimated ± 10.5 pp, low confidence
GLM-4.63.1%19.9%estimated ± 10.5 pp, low confidence
GLM-5 (Reasoning)23.4%19.9%estimated ± 10.5 pp, low confidence
GPT-5.1-Codex13.1%19.9%estimated ± 10.5 pp, low confidence
GPT-5.1-Codex-Max22.2%19.9%estimated ± 10.5 pp, low confidence
GPT-5 mini14.2%19.9%estimated ± 10.5 pp, low confidence
Grok 4.1 Fast (Reasoning)1.2%19.9%estimated ± 10.5 pp, low confidence
Grok 4 Fast (Reasoning)0.0%19.9%estimated ± 10.5 pp, low confidence
MiniMax M2.514.9%19.9%estimated ± 10.5 pp, low confidence
Qwen3.5 Plus15.7%19.9%estimated ± 10.5 pp, low confidence
Qwen3 Max3.5%19.9%estimated ± 10.5 pp, low confidence