benchgap
Calibration

SWE-bench Pro → CursorBench 3.2

CursorBench 3.2 is estimated from SWE-bench Pro with a Hill curve fitted on 12 models measured on both: y = 0.0000 + (0.7496 − 0.0000)·x^6.00 / (0.48799^6.00 + x^6.00), R² = 0.81, cross-validated error 3.6 pp. It is used for 55 estimates.

Estimated modelSWE-bench ProCursorBench 3.2Source
Atria Dawn Preview59.6%57.6%estimated ± 3.6 pp, high confidence
Claude Mythos 580.3%71.4%estimated ± 3.6 pp, high confidence
Claude Opus 4.557.1%53.9%estimated ± 3.6 pp, high confidence
Claude Opus 4.7 (Adaptive)64.3%62.9%estimated ± 3.6 pp, high confidence
DeepSeek V4 Pro 081355.4%51.1%estimated ± 3.6 pp, high confidence
dots3-note Preview61.0%59.4%estimated ± 3.6 pp, high confidence
Gemini 3.5 Flash-Lite54.2%48.9%estimated ± 3.6 pp, medium confidence
GLM-555.1%50.6%estimated ± 3.6 pp, high confidence
GLM-5.158.4%55.9%estimated ± 3.6 pp, high confidence
GPT-5.255.6%51.4%estimated ± 3.6 pp, high confidence
GPT-5.3 Codex56.8%53.5%estimated ± 3.6 pp, high confidence
GPT-5.457.7%54.9%estimated ± 3.6 pp, high confidence
Granite 4.2 30B33.3%6.9%estimated ± 3.6 pp, medium confidence
Granite 4.2 8B19.1%0.3%estimated ± 3.6 pp, medium confidence
Grok 4.2051.8%44.1%estimated ± 3.6 pp, medium confidence
Hy4 preview65.7%64.2%estimated ± 3.6 pp, high confidence
Inkling54.3%49.1%estimated ± 3.6 pp, medium confidence
Inkling-Small55.9%52.0%estimated ± 3.6 pp, high confidence
Kimi K2.658.6%56.2%estimated ± 3.6 pp, high confidence
Kimi K2.550.7%41.8%estimated ± 3.6 pp, medium confidence
Laguna M.149.2%38.4%estimated ± 3.6 pp, medium confidence
Laguna S 2.159.4%57.3%estimated ± 3.6 pp, high confidence
Laguna XS.246.3%31.6%estimated ± 3.6 pp, medium confidence
Laguna XS 2.147.6%34.7%estimated ± 3.6 pp, medium confidence
Ling 3.0 Flash56.6%53.1%estimated ± 3.6 pp, high confidence
LLaDA2.2-flash30.1%3.9%estimated ± 3.6 pp, medium confidence
LongCat-Flash-Lite-Sparse40.6%18.7%estimated ± 3.6 pp, medium confidence
MAI-Thinking-152.8%46.2%estimated ± 3.6 pp, medium confidence
MiMo-V2.556.1%52.3%estimated ± 3.6 pp, high confidence
MiMo-V2.5-Pro57.2%54.1%estimated ± 3.6 pp, high confidence
MiniCPM5-2B14.4%0.0%estimated ± 3.6 pp, medium confidence
MiniMax M2.756.2%52.5%estimated ± 3.6 pp, high confidence
MiniMax M359.0%56.8%estimated ± 3.6 pp, high confidence
Muse Glimmer 30B51.2%42.8%estimated ± 3.6 pp, medium confidence
Muse Spark52.4%45.4%estimated ± 3.6 pp, medium confidence
Muse Spark 1.161.5%60.0%estimated ± 3.6 pp, high confidence
Ornith-1.0-35B50.4%41.1%estimated ± 3.6 pp, medium confidence
Ornith-1.0-397B62.2%60.8%estimated ± 3.6 pp, high confidence
Ornith-1.0-9B42.9%23.7%estimated ± 3.6 pp, medium confidence
Ornith-1.5-35B-A3B59.6%57.6%estimated ± 3.6 pp, high confidence
Ornith-1.5-397B65.1%63.7%estimated ± 3.6 pp, high confidence
Ornith-1.5-9B47.5%34.5%estimated ± 3.6 pp, medium confidence
Qwen3.5 397B50.9%42.2%estimated ± 3.6 pp, medium confidence
Qwen3.6-27B53.5%47.6%estimated ± 3.6 pp, medium confidence
Qwen3.6-35B-A3B49.5%39.1%estimated ± 3.6 pp, medium confidence
Qwen 3.6 Max (preview)57.3%54.3%estimated ± 3.6 pp, high confidence
Qwen3.6 Plus56.6%53.1%estimated ± 3.6 pp, high confidence
Qwen3.7 Max60.6%58.9%estimated ± 3.6 pp, high confidence
Qwen3.7 Plus57.6%54.7%estimated ± 3.6 pp, high confidence
Qwen3.8-Flash-Next62.5%61.1%estimated ± 3.6 pp, high confidence
Qwen3.8-Omni-Flash63.3%62.0%estimated ± 3.6 pp, high confidence
Beam65.5%64.0%estimated ± 3.6 pp, high confidence
Sakana Fugu59.0%56.8%estimated ± 3.6 pp, high confidence
Sakana Fugu-Ultra73.7%69.1%estimated ± 3.6 pp, high confidence
Step 3.7 Flash56.3%52.6%estimated ± 3.6 pp, high confidence