benchgap
Calibration

Vals SWE-bench → NL2Repo

NL2Repo is estimated from Vals SWE-bench with a Hill curve fitted on 13 models measured on both: y = 0.3602 + (1.2000 − 0.3602)·x^6.00 / (1.12009^6.00 + x^6.00), R² = 0.76, cross-validated error 4.7 pp. It is used for 44 estimates.

Estimated modelVals SWE-benchNL2RepoSource
Claude Fable 595.0%58.8%estimated ± 4.7 pp, high confidence
Claude Haiku 4.566.6%39.6%estimated ± 4.7 pp, medium confidence
Claude Opus 4.782.0%47.2%estimated ± 4.7 pp, high confidence
Claude Opus 4.888.6%52.5%estimated ± 4.7 pp, high confidence
Claude Opus 597.0%60.9%estimated ± 4.7 pp, medium confidence
Claude Sonnet 4.677.4%44.3%estimated ± 4.7 pp, high confidence
Claude Sonnet 579.6%45.6%estimated ± 4.7 pp, high confidence
Composer 2.579.6%45.6%estimated ± 4.7 pp, high confidence
Gemini 2.5 Pro54.4%37.1%estimated ± 4.7 pp, medium confidence
Gemini 3.1 Flash-Lite62.8%38.5%estimated ± 4.7 pp, medium confidence
Gemini 3.1 Pro78.8%45.1%estimated ± 4.7 pp, high confidence
Gemini 3.5 Flash78.8%45.1%estimated ± 4.7 pp, high confidence
Gemini 3.5 Flash-Lite75.0%43.0%estimated ± 4.7 pp, high confidence
Gemini 3.6 Flash79.6%45.6%estimated ± 4.7 pp, high confidence
Gemini 3.7 Flash80.8%46.4%estimated ± 4.7 pp, high confidence
Gemini 3.8 Flash80.0%45.9%estimated ± 4.7 pp, high confidence
Gemini 3 Flash75.0%43.0%estimated ± 4.7 pp, high confidence
GLM-4.769.4%40.5%estimated ± 4.7 pp, high confidence
GPT-5.2-Codex72.4%41.7%estimated ± 4.7 pp, high confidence
GPT-5.3 Codex78.0%44.6%estimated ± 4.7 pp, high confidence
GPT-5.4 mini73.0%42.0%estimated ± 4.7 pp, high confidence
GPT-5.4 nano69.8%40.7%estimated ± 4.7 pp, high confidence
GPT-5.582.6%47.7%estimated ± 4.7 pp, high confidence
GPT-5.6 Luna93.0%56.7%estimated ± 4.7 pp, high confidence
GPT-5.6 Sol96.2%60.1%estimated ± 4.7 pp, high confidence
GPT-5.6 Terra95.4%59.2%estimated ± 4.7 pp, high confidence
Grok 4.2072.2%41.6%estimated ± 4.7 pp, high confidence
Grok 4.371.4%41.3%estimated ± 4.7 pp, high confidence
Grok 4.586.6%50.8%estimated ± 4.7 pp, high confidence
Grok 4.695.6%59.4%estimated ± 4.7 pp, high confidence
Inkling77.6%44.4%estimated ± 4.7 pp, high confidence
Inkling-Small82.2%47.4%estimated ± 4.7 pp, high confidence
Kimi K2.7 Code78.2%44.7%estimated ± 4.7 pp, high confidence
Kimi K393.4%57.1%estimated ± 4.7 pp, high confidence
Laguna M.157.6%37.5%estimated ± 4.7 pp, medium confidence
Laguna XS.255.2%37.2%estimated ± 4.7 pp, medium confidence
Ling 3.0 Flash65.2%39.2%estimated ± 4.7 pp, medium confidence
MiMo-V2.571.0%41.1%estimated ± 4.7 pp, high confidence
MiMo-V2.5-Pro74.0%42.5%estimated ± 4.7 pp, high confidence
Mistral Medium 3.5 128B66.4%39.5%estimated ± 4.7 pp, medium confidence
Muse Spark74.4%42.7%estimated ± 4.7 pp, high confidence
Muse Spark 1.182.0%47.2%estimated ± 4.7 pp, high confidence
Muse Spark 1.286.6%50.8%estimated ± 4.7 pp, high confidence
Qwen3.5 Flash64.4%38.9%estimated ± 4.7 pp, medium confidence