benchgap
Calibration

SWE-bench Verified → Vals SWE-bench

Vals SWE-bench is estimated from SWE-bench Verified with a linear curve fitted on 24 models measured on both: y = 1.3755·x + -0.3368, R² = 0.69, cross-validated error 6.8 pp. It is used for 24 estimates.

Estimated modelSWE-bench VerifiedVals SWE-benchSource
BTL-478.4%74.2%estimated ± 6.8 pp, medium confidence
Claude 3.5 Sonnet49.0%33.7%estimated ± 6.8 pp, low confidence
Claude 4.1 Opus74.5%68.8%estimated ± 6.8 pp, medium confidence
Claude 4 Sonnet72.7%66.3%estimated ± 6.8 pp, medium confidence
Claude Mythos 595.5%97.7%estimated ± 6.8 pp, medium confidence
Claude Sonnet 4.577.2%72.5%estimated ± 6.8 pp, medium confidence
Ember-192.2%93.1%estimated ± 6.8 pp, medium confidence
GPT-4.154.6%41.4%estimated ± 6.8 pp, low confidence
Grok Code Fast 170.8%63.7%estimated ± 6.8 pp, medium confidence
Laguna XS 2.170.9%63.8%estimated ± 6.8 pp, medium confidence
LLaDA2.2-flash49.3%34.1%estimated ± 6.8 pp, low confidence
LongCat-Flash-Lite-Sparse68.2%60.1%estimated ± 6.8 pp, medium confidence
MAI-Code-1.1-Flash72.6%66.2%estimated ± 6.8 pp, medium confidence
MAI-Thinking-173.5%67.4%estimated ± 6.8 pp, medium confidence
MiMo-V2-Omni74.8%69.2%estimated ± 6.8 pp, medium confidence
MiMo-V2-Pro78.0%73.6%estimated ± 6.8 pp, medium confidence
o3-mini49.3%34.1%estimated ± 6.8 pp, low confidence
Qwen3.5-27B72.4%65.9%estimated ± 6.8 pp, medium confidence
Qwen3.5-35B-A3B69.2%61.5%estimated ± 6.8 pp, medium confidence
Qwen3.5 397B76.2%71.1%estimated ± 6.8 pp, medium confidence
Solar Open 270.4%63.2%estimated ± 6.8 pp, medium confidence
Solar Pro 470.6%63.4%estimated ± 6.8 pp, medium confidence
Ternary Bonsai 2 27B60.8%50.0%estimated ± 6.8 pp, low confidence
ZAYA1-74B-Preview53.2%39.5%estimated ± 6.8 pp, low confidence