benchgap
Calibration

SWE-bench Verified → LiveCodeBench v6

LiveCodeBench v6 is estimated from SWE-bench Verified with a Hill curve fitted on 15 models measured on both: y = 0.5362 + (0.8783 − 0.5362)·x^6.00 / (0.48357^6.00 + x^6.00), R² = 0.42, cross-validated error 8.1 pp. It is used for 15 estimates.

Estimated modelSWE-bench VerifiedLiveCodeBench v6Source
Claude 3.5 Sonnet49.0%71.4%estimated ± 8.1 pp, low confidence
Claude 4.1 Opus74.5%85.4%estimated ± 8.1 pp, low confidence
Claude 4 Sonnet72.7%85.1%estimated ± 8.1 pp, low confidence
Claude Haiku 4.573.3%85.2%estimated ± 8.1 pp, low confidence
Claude Sonnet 4.577.2%85.9%estimated ± 8.1 pp, low confidence
Claude Sonnet 4.679.6%86.2%estimated ± 8.1 pp, low confidence
Ember-192.2%87.1%estimated ± 8.1 pp, low confidence
GPT-4.154.6%76.7%estimated ± 8.1 pp, low confidence
Grok Code Fast 170.8%84.7%estimated ± 8.1 pp, low confidence
MAI-Code-1.1-Flash72.6%85.1%estimated ± 8.1 pp, low confidence
MiMo-V2-Omni74.8%85.5%estimated ± 8.1 pp, low confidence
MiMo-V2-Pro78.0%86.0%estimated ± 8.1 pp, low confidence
o3-mini49.3%71.7%estimated ± 8.1 pp, low confidence
Qwen3.5-27B72.4%85.0%estimated ± 8.1 pp, low confidence
Qwen3.5-35B-A3B69.2%84.3%estimated ± 8.1 pp, low confidence