benchgap
Calibration

AA-SciCode → CursorBench 3.1

CursorBench 3.1 is estimated from AA-SciCode with a linear curve fitted on 5 models measured on both: y = 2.4655·x + -0.7927, R² = 0.91, cross-validated error 3.4 pp. It is used for 100 estimates.

Estimated modelAA-SciCodeCursorBench 3.1Source
A.X K241.0%21.8%estimated ± 3.4 pp, low confidence
Apodex 1.145.5%32.9%estimated ± 3.4 pp, low confidence
Apodex 1.1 Mini45.5%32.9%estimated ± 3.4 pp, low confidence
Celeris-121.6%0.0%estimated ± 3.4 pp, low confidence
Claude Fable 5.163.1%76.3%estimated ± 3.4 pp, low confidence
Claude Haiku 5.555.0%56.3%estimated ± 3.4 pp, medium confidence
Claude Opus 556.4%59.8%estimated ± 3.4 pp, medium confidence
Claude Opus 5.566.9%85.7%estimated ± 3.4 pp, low confidence
Claude Sonnet 554.3%54.6%estimated ± 3.4 pp, medium confidence
Claude Sonnet 5.561.0%71.1%estimated ± 3.4 pp, medium confidence
Command A+38.5%15.6%estimated ± 3.4 pp, low confidence
DeepSeek V335.8%9.0%estimated ± 3.4 pp, low confidence
DeepSeek V3 032439.0%16.9%estimated ± 3.4 pp, low confidence
DeepSeek V4.1 Flash51.9%48.7%estimated ± 3.4 pp, medium confidence
DeepSeek V4 Flash 073150.3%44.7%estimated ± 3.4 pp, low confidence
DeepSeek V4 Pro 081351.0%46.5%estimated ± 3.4 pp, low confidence
Gemini 2.5 Pro46.3%34.9%estimated ± 3.4 pp, low confidence
Gemini 3.1 Pro58.7%65.5%estimated ± 3.4 pp, medium confidence
Gemini 3.5 Flash-Lite41.3%22.6%estimated ± 3.4 pp, low confidence
Gemini 3.6 Flash53.4%52.4%estimated ± 3.4 pp, medium confidence
Gemini 3.7 Flash57.2%61.8%estimated ± 3.4 pp, medium confidence
Gemini 3.8 Flash56.6%60.3%estimated ± 3.4 pp, medium confidence
Gemini 4 Argon61.8%73.1%estimated ± 3.4 pp, low confidence
Gemma 3 27B23.3%0.0%estimated ± 3.4 pp, low confidence
Gemma 4 26B A4B40.0%19.3%estimated ± 3.4 pp, low confidence
Gemma 4 31B45.5%32.9%estimated ± 3.4 pp, low confidence
Gemma 4 E4B24.4%0.0%estimated ± 3.4 pp, low confidence
GLM-5.144.8%31.2%estimated ± 3.4 pp, low confidence
GLM-5.251.2%47.0%estimated ± 3.4 pp, low confidence
GLM-5.359.0%66.2%estimated ± 3.4 pp, medium confidence
GLM-5.3-Flash51.6%47.9%estimated ± 3.4 pp, medium confidence
GPT-5.4 mini52.1%49.2%estimated ± 3.4 pp, medium confidence
GPT-5.4 nano47.2%37.1%estimated ± 3.4 pp, low confidence
GPT-5.6 Luna53.6%52.9%estimated ± 3.4 pp, medium confidence
GPT-5.6 Sol57.1%61.5%estimated ± 3.4 pp, medium confidence
GPT-5.6 Terra55.0%56.3%estimated ± 3.4 pp, medium confidence
GPT-6.1 Sol54.2%54.4%estimated ± 3.4 pp, medium confidence
GPT-6 Astra56.5%60.0%estimated ± 3.4 pp, medium confidence
GPT-6 Luna54.6%55.3%estimated ± 3.4 pp, medium confidence
GPT-6 Sol57.6%62.7%estimated ± 3.4 pp, medium confidence
GPT-OSS 120B34.0%4.6%estimated ± 3.4 pp, low confidence
GPT-OSS 20B38.9%16.6%estimated ± 3.4 pp, low confidence
Granite 4.2 30B37.8%13.9%estimated ± 3.4 pp, low confidence
Granite 4.2 3B25.3%0.0%estimated ± 3.4 pp, low confidence
Granite 4.2 8B31.5%0.0%estimated ± 3.4 pp, low confidence
Grok 4.348.3%39.8%estimated ± 3.4 pp, low confidence
Grok 4.555.0%56.3%estimated ± 3.4 pp, medium confidence
Grok 4.656.5%60.0%estimated ± 3.4 pp, medium confidence
Grok 4.757.4%62.2%estimated ± 3.4 pp, medium confidence
Hy348.6%40.6%estimated ± 3.4 pp, low confidence
Hy3 Preview48.6%40.6%estimated ± 3.4 pp, low confidence
Inkling47.0%36.6%estimated ± 3.4 pp, low confidence
Inkling-Small49.7%43.3%estimated ± 3.4 pp, low confidence
K-EXAONE 2.042.0%24.3%estimated ± 3.4 pp, low confidence
Kimi K2.7 Code47.8%38.6%estimated ± 3.4 pp, low confidence
Kimi K359.5%67.4%estimated ± 3.4 pp, medium confidence
LFM2.5-2.6B14.4%0.0%estimated ± 3.4 pp, low confidence
Ling 3.0 Flash42.0%24.3%estimated ± 3.4 pp, low confidence
Ling 3.0 Flash FP842.0%24.3%estimated ± 3.4 pp, low confidence
Ling 3.0 Flash VL44.2%29.7%estimated ± 3.4 pp, low confidence
Ling 3.0 Tiny24.2%0.0%estimated ± 3.4 pp, low confidence
Ling 3.1 Flash54.1%54.1%estimated ± 3.4 pp, medium confidence
Llama 4 Maverick31.7%0.0%estimated ± 3.4 pp, low confidence
Llama 4 Scout21.3%0.0%estimated ± 3.4 pp, low confidence
Mercury 2.538.5%15.6%estimated ± 3.4 pp, low confidence
MiMo-V2.5-Pro50.6%45.5%estimated ± 3.4 pp, low confidence
MiMo-V2.6-Flash51.3%47.2%estimated ± 3.4 pp, low confidence
MiMo-V2.6-Pro60.9%70.9%estimated ± 3.4 pp, medium confidence
MiniCPM5-2B26.3%0.0%estimated ± 3.4 pp, low confidence
MiniMax M2.750.1%44.3%estimated ± 3.4 pp, low confidence
MiniMax M347.1%36.9%estimated ± 3.4 pp, low confidence
Mistral Large 336.6%11.0%estimated ± 3.4 pp, low confidence
Mistral Large 454.2%54.4%estimated ± 3.4 pp, medium confidence
Mistral Medium 3.5 128B40.2%19.8%estimated ± 3.4 pp, low confidence
Mistral Small 438.8%16.4%estimated ± 3.4 pp, low confidence
Mistral Small 4 (Reasoning)38.8%16.4%estimated ± 3.4 pp, low confidence
Muse Glimmer 30B44.9%31.4%estimated ± 3.4 pp, low confidence
Muse Spark 1.158.8%65.7%estimated ± 3.4 pp, medium confidence
Muse Spark 1.257.4%62.2%estimated ± 3.4 pp, medium confidence
Muse Spark 1.358.8%65.7%estimated ± 3.4 pp, medium confidence
Nemotron 3.5 Lightning 30B A3B NVFP432.1%0.0%estimated ± 3.4 pp, low confidence
Nemotron 3 Nano 30B30.6%0.0%estimated ± 3.4 pp, low confidence
Nemotron 3 Super 100B36.2%10.0%estimated ± 3.4 pp, low confidence
Nemotron 3 Ultra40.3%20.1%estimated ± 3.4 pp, low confidence
North Mini Code38.8%16.4%estimated ± 3.4 pp, low confidence
Quasar 438B48.1%39.3%estimated ± 3.4 pp, low confidence
Qwen3.5-122B-A10B39.7%18.6%estimated ± 3.4 pp, low confidence
Qwen3.6-27B42.8%26.3%estimated ± 3.4 pp, low confidence
Qwen3.6-35B-A3B36.6%11.0%estimated ± 3.4 pp, low confidence
Qwen3.7 Max49.5%42.8%estimated ± 3.4 pp, low confidence
Qwen3.7 Plus46.1%34.4%estimated ± 3.4 pp, low confidence
Qwen3.8-27B46.6%35.6%estimated ± 3.4 pp, low confidence
Qwen3.8-Flash-Next50.6%45.5%estimated ± 3.4 pp, low confidence
Qwen3.8 Max Preview52.1%49.2%estimated ± 3.4 pp, medium confidence
Solar Pro 325.5%0.0%estimated ± 3.4 pp, low confidence
Solar Pro 444.6%30.7%estimated ± 3.4 pp, low confidence
Step 3.7 Flash43.9%29.0%estimated ± 3.4 pp, low confidence
Step 5 Preview58.9%65.9%estimated ± 3.4 pp, medium confidence
Trinity-Large-Preview40.6%20.8%estimated ± 3.4 pp, low confidence
Trinity-Large-Thinking40.6%20.8%estimated ± 3.4 pp, low confidence