benchgap
Calibration

DeepSWE → SWE-bench Verified

SWE-bench Verified is estimated from DeepSWE with a Hill curve fitted on 7 models measured on both: y = 0.7866 + (1.1365 − 0.7866)·x^6.00 / (0.78061^6.00 + x^6.00), R² = 0.67, cross-validated error 5.2 pp. It is used for 10 estimates.

Estimated modelDeepSWESWE-bench VerifiedSource
GPT-6.1 Sol71.9%91.9%estimated ± 5.2 pp, low confidence
GPT-6 Luna66.6%88.4%estimated ± 5.2 pp, low confidence
GPT-6 Sol68.8%89.8%estimated ± 5.2 pp, low confidence
Grok 4.771.0%91.3%estimated ± 5.2 pp, low confidence
MiMo-V2.6-Flash67.9%89.2%estimated ± 5.2 pp, low confidence
MiMo-V2.6-Pro71.9%91.9%estimated ± 5.2 pp, low confidence
Muse Spark 1.375.4%94.3%estimated ± 5.2 pp, low confidence
Pareto 26.10 Preview69.9%90.6%estimated ± 5.2 pp, low confidence
Pareto 26.974.0%93.4%estimated ± 5.2 pp, low confidence
Step 5 Preview67.7%89.1%estimated ± 5.2 pp, low confidence