benchgap
Calibration

SWE-bench Verified → AA-SciCode

AA-SciCode is estimated from SWE-bench Verified with a Michaelis–Menten curve fitted on 28 models measured on both: y = 2.0000·x / (2.51760 + x), R² = 0.77, cross-validated error 3.9 pp. It is used for 14 estimates.

Estimated modelSWE-bench VerifiedAA-SciCodeSource
Claude 3.5 Sonnet49.0%32.6%estimated ± 3.9 pp, high confidence
Claude 4.1 Opus74.5%45.7%estimated ± 3.9 pp, high confidence
Claude 4 Sonnet72.7%44.8%estimated ± 3.9 pp, high confidence
Claude Haiku 4.573.3%45.1%estimated ± 3.9 pp, high confidence
Claude Sonnet 4.577.2%46.9%estimated ± 3.9 pp, high confidence
Ember-192.2%53.6%estimated ± 3.9 pp, high confidence
GPT-4.154.6%35.6%estimated ± 3.9 pp, high confidence
Grok Code Fast 170.8%43.9%estimated ± 3.9 pp, high confidence
MAI-Code-1.1-Flash72.6%44.8%estimated ± 3.9 pp, high confidence
MiMo-V2-Omni74.8%45.8%estimated ± 3.9 pp, high confidence
MiMo-V2-Pro78.0%47.3%estimated ± 3.9 pp, high confidence
o3-mini49.3%32.8%estimated ± 3.9 pp, high confidence
Qwen3.5-27B72.4%44.7%estimated ± 3.9 pp, high confidence
Qwen3.5-35B-A3B69.2%43.1%estimated ± 3.9 pp, high confidence