benchgap
Calibration

SWE-bench Verified → NL2Repo

NL2Repo is estimated from SWE-bench Verified with a Hill curve fitted on 15 models measured on both: y = 0.0131 + (1.1525 − 0.0131)·x^6.00 / (0.85375^6.00 + x^6.00), R² = 0.70, cross-validated error 6.1 pp. It is used for 23 estimates.

Estimated modelSWE-bench VerifiedNL2RepoSource
Claude 3.5 Sonnet49.0%5.2%estimated ± 6.1 pp, low confidence
Claude 4.1 Opus74.5%36.2%estimated ± 6.1 pp, medium confidence
Claude 4 Sonnet72.7%32.8%estimated ± 6.1 pp, medium confidence
Claude Mythos 595.5%76.7%estimated ± 6.1 pp, low confidence
Claude Opus 4.680.8%49.0%estimated ± 6.1 pp, medium confidence
Claude Sonnet 4.577.2%41.6%estimated ± 6.1 pp, medium confidence
Ember-192.2%71.2%estimated ± 6.1 pp, low confidence
GLM-577.8%42.8%estimated ± 6.1 pp, medium confidence
GPT-4.154.6%8.6%estimated ± 6.1 pp, low confidence
GPT-5.280.0%47.3%estimated ± 6.1 pp, medium confidence
Grok Code Fast 170.8%29.3%estimated ± 6.1 pp, medium confidence
K-EXAONE 2.068.2%24.8%estimated ± 6.1 pp, low confidence
Laguna XS 2.170.9%29.5%estimated ± 6.1 pp, medium confidence
LLaDA2.2-flash49.3%5.4%estimated ± 6.1 pp, low confidence
LongCat-Flash-Lite-Sparse68.2%24.8%estimated ± 6.1 pp, low confidence
MAI-Code-1.1-Flash72.6%32.6%estimated ± 6.1 pp, medium confidence
MiMo-V2-Omni74.8%36.8%estimated ± 6.1 pp, medium confidence
MiMo-V2-Pro78.0%43.2%estimated ± 6.1 pp, medium confidence
o3-mini49.3%5.4%estimated ± 6.1 pp, low confidence
Qwen3.5-27B72.4%32.2%estimated ± 6.1 pp, medium confidence
Qwen3.5-35B-A3B69.2%26.5%estimated ± 6.1 pp, low confidence
Beam80.9%49.2%estimated ± 6.1 pp, medium confidence
Solar Pro 470.6%28.9%estimated ± 6.1 pp, medium confidence