benchgap
Calibration

Vals SWE-bench → PostTrainBench v1.1

PostTrainBench v1.1 is estimated from Vals SWE-bench with a inverse Michaelis–Menten curve fitted on 9 models measured on both: y = 0.38080·(x − 0.0000) / (0.0000 + 2.0000 − x), R² = 0.61, cross-validated error 3.3 pp. It is used for 8 estimates.

Estimated modelVals SWE-benchPostTrainBench v1.1Source
Composer 2.579.6%25.2%estimated ± 3.3 pp, high confidence
Gemini 2.5 Pro54.4%14.2%estimated ± 3.3 pp, medium confidence
GPT-5.6 Luna93.0%33.1%estimated ± 3.3 pp, high confidence
Mistral Medium 3.5 128B66.4%18.9%estimated ± 3.3 pp, medium confidence
Muse Spark74.4%22.6%estimated ± 3.3 pp, medium confidence
Muse Spark 1.286.6%29.1%estimated ± 3.3 pp, high confidence
Qwen3.6-27B70.0%20.5%estimated ± 3.3 pp, medium confidence
Qwen 3.6 Max (preview)72.8%21.8%estimated ± 3.3 pp, medium confidence