benchgap
Calibration

SWE-bench Pro → SWE-bench Verified

SWE-bench Verified is estimated from SWE-bench Pro with a Hill curve fitted on 44 models measured on both: y = 0.4294 + (1.1152 − 0.4294)·x^2.67 / (0.53491^2.67 + x^2.67), R² = 0.95, cross-validated error 2.5 pp. It is used for 17 estimates.

Estimated modelSWE-bench ProSWE-bench VerifiedSource
Atria Dawn Preview59.6%82.2%estimated ± 2.5 pp, high confidence
Claude Fable 5.181.2%94.6%estimated ± 2.5 pp, medium confidence
Gemini 3.5 Flash55.1%78.6%estimated ± 2.5 pp, high confidence
Gemini 3.5 Flash-Lite54.2%77.8%estimated ± 2.5 pp, high confidence
GPT-5.457.7%80.7%estimated ± 2.5 pp, high confidence
GPT-5.6 Luna62.7%84.4%estimated ± 2.5 pp, high confidence
GPT-5.6 Sol64.6%85.7%estimated ± 2.5 pp, high confidence
GPT-5.6 Terra63.4%84.9%estimated ± 2.5 pp, high confidence
Grok 4.564.7%85.8%estimated ± 2.5 pp, high confidence
Laguna S 2.159.4%82.0%estimated ± 2.5 pp, high confidence
Ling 3.0 Flash56.6%79.8%estimated ± 2.5 pp, high confidence
MiMo-V2.556.1%79.4%estimated ± 2.5 pp, high confidence
MiMo-V2.5-Pro57.2%80.3%estimated ± 2.5 pp, high confidence
Muse Spark 1.161.5%83.5%estimated ± 2.5 pp, high confidence
Sakana Fugu59.0%81.7%estimated ± 2.5 pp, high confidence
Sakana Fugu-Ultra73.7%91.1%estimated ± 2.5 pp, high confidence
Step 3.7 Flash56.3%79.6%estimated ± 2.5 pp, high confidence