benchgap
Calibration

SWE-bench Pro → FrontierCode 1.1 Extended

FrontierCode 1.1 Extended is estimated from SWE-bench Pro with a Michaelis–Menten curve fitted on 6 models measured on both: y = 0.8660·x / (0.32783 + x), R² = 0.60, cross-validated error 3.0 pp. It is used for 35 estimates.

Estimated modelSWE-bench ProFrontierCode 1.1 ExtendedSource
Atria Dawn Preview59.6%55.9%estimated ± 3.0 pp, low confidence
Claude Mythos 580.3%61.5%estimated ± 3.0 pp, medium confidence
Claude Opus 4.557.1%55.0%estimated ± 3.0 pp, low confidence
Claude Opus 4.653.4%53.7%estimated ± 3.0 pp, low confidence
Claude Opus 4.7 (Adaptive)64.3%57.4%estimated ± 3.0 pp, medium confidence
dots3-note Preview61.0%56.3%estimated ± 3.0 pp, low confidence
GLM-555.1%54.3%estimated ± 3.0 pp, low confidence
GPT-5.255.6%54.5%estimated ± 3.0 pp, low confidence
GPT-5.457.7%55.2%estimated ± 3.0 pp, low confidence
Granite 4.2 30B33.3%43.6%estimated ± 3.0 pp, low confidence
Granite 4.2 8B19.1%31.9%estimated ± 3.0 pp, low confidence
Hy4 preview65.7%57.8%estimated ± 3.0 pp, medium confidence
Kimi K2.550.7%52.6%estimated ± 3.0 pp, low confidence
Laguna S 2.159.4%55.8%estimated ± 3.0 pp, low confidence
Laguna XS 2.147.6%51.3%estimated ± 3.0 pp, low confidence
LLaDA2.2-flash30.1%41.5%estimated ± 3.0 pp, low confidence
LongCat-Flash-Lite-Sparse40.6%47.9%estimated ± 3.0 pp, low confidence
MAI-Thinking-152.8%53.4%estimated ± 3.0 pp, low confidence
MiniCPM5-2B14.4%26.4%estimated ± 3.0 pp, low confidence
Muse Glimmer 30B51.2%52.8%estimated ± 3.0 pp, low confidence
Ornith-1.0-35B50.4%52.5%estimated ± 3.0 pp, low confidence
Ornith-1.0-397B62.2%56.7%estimated ± 3.0 pp, low confidence
Ornith-1.0-9B42.9%49.1%estimated ± 3.0 pp, low confidence
Ornith-1.5-35B-A3B59.6%55.9%estimated ± 3.0 pp, low confidence
Ornith-1.5-397B65.1%57.6%estimated ± 3.0 pp, medium confidence
Ornith-1.5-9B47.5%51.2%estimated ± 3.0 pp, low confidence
Qwen3.5 397B50.9%52.7%estimated ± 3.0 pp, low confidence
Qwen3.6-35B-A3B49.5%52.1%estimated ± 3.0 pp, low confidence
Qwen3.7 Plus57.6%55.2%estimated ± 3.0 pp, low confidence
Qwen3.8-Flash-Next62.5%56.8%estimated ± 3.0 pp, low confidence
Qwen3.8-Omni-Flash63.3%57.1%estimated ± 3.0 pp, medium confidence
Beam65.5%57.7%estimated ± 3.0 pp, medium confidence
Sakana Fugu59.0%55.7%estimated ± 3.0 pp, low confidence
Sakana Fugu-Ultra73.7%59.9%estimated ± 3.0 pp, medium confidence
Step 3.7 Flash56.3%54.7%estimated ± 3.0 pp, low confidence