benchgap
Calibration

Vals SWE-bench → NL2Repo

NL2Repo is estimated from Vals SWE-bench with a Hill curve fitted on 14 models measured on both: y = 0.3517 + (1.2000 − 0.3517)·x^6.00 / (1.11211^6.00 + x^6.00), R² = 0.77, cross-validated error 4.7 pp. It is used for 43 estimates.

Estimated modelVals SWE-benchNL2RepoSource
Claude Opus 597.0%61.1%estimated ± 4.7 pp, medium confidence
GPT-5.6 Sol96.2%60.2%estimated ± 4.7 pp, high confidence
Grok 4.695.6%59.6%estimated ± 4.7 pp, high confidence
GPT-5.6 Terra95.4%59.3%estimated ± 4.7 pp, high confidence
Claude Fable 595.0%58.9%estimated ± 4.7 pp, high confidence
Kimi K393.4%57.2%estimated ± 4.7 pp, high confidence
GPT-5.6 Luna93.0%56.8%estimated ± 4.7 pp, high confidence
Claude Opus 4.888.6%52.4%estimated ± 4.7 pp, high confidence
Grok 4.586.6%50.6%estimated ± 4.7 pp, high confidence
Muse Spark 1.286.6%50.6%estimated ± 4.7 pp, high confidence
GPT-5.582.6%47.4%estimated ± 4.7 pp, high confidence
Inkling-Small82.2%47.1%estimated ± 4.7 pp, high confidence
Claude Opus 4.782.0%46.9%estimated ± 4.7 pp, high confidence
Muse Spark 1.182.0%46.9%estimated ± 4.7 pp, high confidence
Gemini 3.7 Flash80.8%46.0%estimated ± 4.7 pp, high confidence
Gemini 3.8 Flash80.0%45.5%estimated ± 4.7 pp, high confidence
Claude Sonnet 579.6%45.2%estimated ± 4.7 pp, high confidence
Composer 2.579.6%45.2%estimated ± 4.7 pp, high confidence
Gemini 3.6 Flash79.6%45.2%estimated ± 4.7 pp, high confidence
Gemini 3.1 Pro78.8%44.7%estimated ± 4.7 pp, high confidence
Gemini 3.5 Flash78.8%44.7%estimated ± 4.7 pp, high confidence
Kimi K2.7 Code78.2%44.3%estimated ± 4.7 pp, high confidence
GPT-5.3 Codex78.0%44.2%estimated ± 4.7 pp, high confidence
Inkling77.6%43.9%estimated ± 4.7 pp, high confidence
Claude Sonnet 4.677.4%43.8%estimated ± 4.7 pp, high confidence
Gemini 3.5 Flash-Lite75.0%42.5%estimated ± 4.7 pp, high confidence
Gemini 3 Flash75.0%42.5%estimated ± 4.7 pp, high confidence
Muse Spark74.4%42.1%estimated ± 4.7 pp, high confidence
MiMo-V2.5-Pro74.0%41.9%estimated ± 4.7 pp, high confidence
GPT-5.4 mini73.0%41.5%estimated ± 4.7 pp, high confidence
GPT-5.2-Codex72.4%41.2%estimated ± 4.7 pp, high confidence
Grok 4.2072.2%41.1%estimated ± 4.7 pp, high confidence
Grok 4.371.4%40.7%estimated ± 4.7 pp, high confidence
MiMo-V2.571.0%40.5%estimated ± 4.7 pp, high confidence
GPT-5.4 nano69.8%40.1%estimated ± 4.7 pp, high confidence
Claude Haiku 4.566.6%38.9%estimated ± 4.7 pp, medium confidence
Mistral Medium 3.5 128B66.4%38.8%estimated ± 4.7 pp, medium confidence
Ling 3.0 Flash65.2%38.5%estimated ± 4.7 pp, medium confidence
Qwen3.5 Flash64.4%38.2%estimated ± 4.7 pp, medium confidence
Gemini 3.1 Flash-Lite62.8%37.8%estimated ± 4.7 pp, medium confidence
Laguna M.157.6%36.8%estimated ± 4.7 pp, medium confidence
Laguna XS.255.2%36.4%estimated ± 4.7 pp, medium confidence
Gemini 2.5 Pro54.4%36.3%estimated ± 4.7 pp, medium confidence