benchgap
Calibration

DeepSWE → PostTrainBench v1.1

PostTrainBench v1.1 is estimated from DeepSWE with a linear curve fitted on 8 models measured on both: y = 0.9562·x + -0.2819, R² = 0.76, cross-validated error 4.8 pp. It is used for 12 estimates.

Estimated modelDeepSWEPostTrainBench v1.1Source
DeepSeek V4.1 Flash74.2%42.8%estimated ± 4.8 pp, high confidence
Ember-175.2%43.7%estimated ± 4.8 pp, high confidence
GPT-6.1 Sol71.9%40.6%estimated ± 4.8 pp, high confidence
GPT-6 Luna66.6%35.5%estimated ± 4.8 pp, high confidence
GPT-6 Sol68.8%37.6%estimated ± 4.8 pp, high confidence
MiMo-V2.6-Flash67.9%36.7%estimated ± 4.8 pp, high confidence
MiMo-V2.6-Pro71.9%40.6%estimated ± 4.8 pp, high confidence
Muse Spark 1.375.4%43.9%estimated ± 4.8 pp, high confidence
Pareto 26.10 Preview69.9%38.7%estimated ± 4.8 pp, high confidence
Pareto 26.974.0%42.6%estimated ± 4.8 pp, high confidence
Step 5 Preview67.7%36.5%estimated ± 4.8 pp, high confidence
SWE-273.0%41.6%estimated ± 4.8 pp, high confidence