benchgap
Calibration

SWE-bench Verified → Vals LiveCodeBench

Vals LiveCodeBench is estimated from SWE-bench Verified with a Hill curve fitted on 21 models measured on both: y = 0.0000 + (1.0017 − 0.0000)·x^6.00 / (0.61912^6.00 + x^6.00), R² = 0.39, cross-validated error 9.9 pp. It is used for 25 estimates.

Estimated modelSWE-bench VerifiedVals LiveCodeBenchSource
Apodex 1.177.7%79.8%estimated ± 9.9 pp, low confidence
BTL-478.4%80.6%estimated ± 9.9 pp, low confidence
Claude 3.5 Sonnet49.0%19.8%estimated ± 9.9 pp, low confidence
Claude 4.1 Opus74.5%75.3%estimated ± 9.9 pp, low confidence
Claude 4 Sonnet72.7%72.5%estimated ± 9.9 pp, low confidence
Claude Sonnet 4.577.2%79.1%estimated ± 9.9 pp, low confidence
DeepSeek V342.0%8.9%estimated ± 9.9 pp, low confidence
Ember-192.2%91.8%estimated ± 9.9 pp, low confidence
Gemini 2.5 Pro63.8%54.6%estimated ± 9.9 pp, low confidence
GPT-4.154.6%32.0%estimated ± 9.9 pp, low confidence
GPT-4.1 mini23.6%0.3%estimated ± 9.9 pp, low confidence
Kimi K2.5 (Reasoning)76.8%78.6%estimated ± 9.9 pp, low confidence
MAI-Code-1.1-Flash72.6%72.3%estimated ± 9.9 pp, low confidence
MiMo-V2-Flash73.4%73.6%estimated ± 9.9 pp, low confidence
MiMo-V2-Omni74.8%75.8%estimated ± 9.9 pp, low confidence
MiMo-V2-Pro78.0%80.1%estimated ± 9.9 pp, low confidence
Mistral Medium 3.5 128B77.6%79.6%estimated ± 9.9 pp, low confidence
o3-mini49.3%20.3%estimated ± 9.9 pp, low confidence
Qwen3.5-122B-A10B72.0%71.3%estimated ± 9.9 pp, low confidence
Qwen3.5-27B72.4%72.0%estimated ± 9.9 pp, low confidence
Qwen3.5-35B-A3B69.2%66.2%estimated ± 9.9 pp, low confidence
Solar Open 270.4%68.5%estimated ± 9.9 pp, low confidence
Solar Pro 470.6%68.9%estimated ± 9.9 pp, low confidence
Ternary Bonsai 2 27B60.8%47.4%estimated ± 9.9 pp, low confidence
ZAYA1-74B-Preview53.2%28.7%estimated ± 9.9 pp, low confidence