benchgap
Calibration

SWE-bench Pro → PostTrainBench v1.1

PostTrainBench v1.1 is estimated from SWE-bench Pro with a inverse Michaelis–Menten curve fitted on 9 models measured on both: y = 0.59838·(x − 0.0000) / (0.0000 + 2.0000 − x), R² = 0.75, cross-validated error 4.7 pp. It is used for 34 estimates.

Estimated modelSWE-bench ProPostTrainBench v1.1Source
Atria Dawn Preview59.6%25.4%estimated ± 4.7 pp, high confidence
Claude Mythos 580.3%40.1%estimated ± 4.7 pp, high confidence
Claude Opus 4.557.1%23.9%estimated ± 4.7 pp, medium confidence
Claude Opus 4.653.4%21.8%estimated ± 4.7 pp, medium confidence
Claude Opus 4.7 (Adaptive)64.3%28.4%estimated ± 4.7 pp, high confidence
dots3-note Preview61.0%26.3%estimated ± 4.7 pp, high confidence
GLM-555.1%22.8%estimated ± 4.7 pp, medium confidence
GPT-5.255.6%23.0%estimated ± 4.7 pp, medium confidence
Granite 4.2 30B33.3%11.9%estimated ± 4.7 pp, medium confidence
Granite 4.2 8B19.1%6.3%estimated ± 4.7 pp, medium confidence
Hy4 preview65.7%29.3%estimated ± 4.7 pp, high confidence
Kimi K2.550.7%20.3%estimated ± 4.7 pp, medium confidence
Laguna S 2.159.4%25.3%estimated ± 4.7 pp, high confidence
Laguna XS 2.147.6%18.7%estimated ± 4.7 pp, medium confidence
LLaDA2.2-flash30.1%10.6%estimated ± 4.7 pp, medium confidence
LongCat-Flash-Lite-Sparse40.6%15.3%estimated ± 4.7 pp, medium confidence
MAI-Thinking-152.8%21.5%estimated ± 4.7 pp, medium confidence
MiniCPM5-2B14.4%4.6%estimated ± 4.7 pp, medium confidence
Muse Glimmer 30B51.2%20.6%estimated ± 4.7 pp, medium confidence
Ornith-1.0-35B50.4%20.2%estimated ± 4.7 pp, medium confidence
Ornith-1.0-397B62.2%27.0%estimated ± 4.7 pp, high confidence
Ornith-1.0-9B42.9%16.3%estimated ± 4.7 pp, medium confidence
Ornith-1.5-35B-A3B59.6%25.4%estimated ± 4.7 pp, high confidence
Ornith-1.5-397B65.1%28.9%estimated ± 4.7 pp, high confidence
Ornith-1.5-9B47.5%18.6%estimated ± 4.7 pp, medium confidence
Qwen3.5 397B50.9%20.4%estimated ± 4.7 pp, medium confidence
Qwen3.6-35B-A3B49.5%19.7%estimated ± 4.7 pp, medium confidence
Qwen3.7 Plus57.6%24.2%estimated ± 4.7 pp, medium confidence
Qwen3.8-Flash-Next62.5%27.2%estimated ± 4.7 pp, high confidence
Qwen3.8-Omni-Flash63.3%27.7%estimated ± 4.7 pp, high confidence
Beam65.5%29.1%estimated ± 4.7 pp, high confidence
Sakana Fugu59.0%25.0%estimated ± 4.7 pp, high confidence
Sakana Fugu-Ultra73.7%34.9%estimated ± 4.7 pp, high confidence
Step 3.7 Flash56.3%23.4%estimated ± 4.7 pp, medium confidence