benchgap
Calibration

AA Coding Index → CursorBench 4.0

CursorBench 4.0 is estimated from AA Coding Index with a inverse Michaelis–Menten curve fitted on 9 models measured on both: y = 0.20585·(x − 0.0000) / (0.0000 + 1.1424 − x), R² = 0.93, cross-validated error 1.6 pp. It is used for 93 estimates.

Estimated modelAA Coding IndexCursorBench 4.0Source
Apodex 1.160.8%23.4%estimated ± 1.6 pp, medium confidence
Apodex 1.1 Mini60.8%23.4%estimated ± 1.6 pp, medium confidence
Celeris-114.4%3.0%estimated ± 1.6 pp, medium confidence
Claude 3 Opus19.5%4.2%estimated ± 1.6 pp, medium confidence
Claude Fable 576.5%41.7%estimated ± 1.6 pp, high confidence
Claude Opus 4.7 (Adaptive)73.6%37.3%estimated ± 1.6 pp, high confidence
Claude Opus 4.874.3%38.2%estimated ± 1.6 pp, high confidence
Command A+27.9%6.6%estimated ± 1.6 pp, medium confidence
DeepSeek V323.0%5.2%estimated ± 1.6 pp, medium confidence
DeepSeek V4 Flash 073169.1%31.5%estimated ± 1.6 pp, medium confidence
DeepSeek V4 Pro 081368.8%31.2%estimated ± 1.6 pp, medium confidence
Gemini 1.5 Pro23.6%5.4%estimated ± 1.6 pp, medium confidence
Gemini 2.5 Pro33.3%8.5%estimated ± 1.6 pp, medium confidence
Gemini 3.1 Pro68.8%31.2%estimated ± 1.6 pp, medium confidence
Gemini 3.5 Flash70.1%32.7%estimated ± 1.6 pp, medium confidence
Gemini 3.5 Flash-Lite49.3%15.6%estimated ± 1.6 pp, medium confidence
Gemini 3.6 Flash69.2%31.7%estimated ± 1.6 pp, medium confidence
Gemini 3.7 Flash76.1%41.1%estimated ± 1.6 pp, high confidence
Gemma 3 27B10.1%2.0%estimated ± 1.6 pp, medium confidence
Gemma 4 12B31.0%7.7%estimated ± 1.6 pp, medium confidence
Gemma 4 26B A4B39.3%10.8%estimated ± 1.6 pp, medium confidence
Gemma 4 31B43.4%12.6%estimated ± 1.6 pp, medium confidence
Gemma 4 E2B7.2%1.4%estimated ± 1.6 pp, medium confidence
Gemma 4 E4B9.4%1.8%estimated ± 1.6 pp, medium confidence
GLM-4.745.3%13.5%estimated ± 1.6 pp, medium confidence
GLM-5.155.8%19.6%estimated ± 1.6 pp, medium confidence
GLM-5.268.8%31.1%estimated ± 1.6 pp, medium confidence
GLM-5.374.8%39.0%estimated ± 1.6 pp, high confidence
GPT-4.1 mini20.2%4.4%estimated ± 1.6 pp, medium confidence
GPT-4.1 nano11.1%2.2%estimated ± 1.6 pp, medium confidence
GPT-4 Turbo21.5%4.8%estimated ± 1.6 pp, medium confidence
GPT-4o mini11.4%2.3%estimated ± 1.6 pp, medium confidence
GPT-5.149.4%15.7%estimated ± 1.6 pp, medium confidence
GPT-5.471.1%33.9%estimated ± 1.6 pp, medium confidence
GPT-5.4 mini56.1%19.8%estimated ± 1.6 pp, medium confidence
GPT-5.4 nano56.1%19.8%estimated ± 1.6 pp, medium confidence
GPT-5.574.9%39.2%estimated ± 1.6 pp, high confidence
GPT-5 (high)37.8%10.2%estimated ± 1.6 pp, medium confidence
GPT-6 Astra76.9%42.5%estimated ± 1.6 pp, high confidence
GPT-OSS 120B30.4%7.5%estimated ± 1.6 pp, medium confidence
GPT-OSS 20B20.7%4.6%estimated ± 1.6 pp, medium confidence
Granite 4.2 8B22.4%5.0%estimated ± 1.6 pp, medium confidence
Grok 4.342.3%12.1%estimated ± 1.6 pp, medium confidence
Grok 4.572.5%35.7%estimated ± 1.6 pp, high confidence
Hy358.8%21.8%estimated ± 1.6 pp, medium confidence
Hy3 Preview58.8%21.8%estimated ± 1.6 pp, medium confidence
Inkling52.1%17.2%estimated ± 1.6 pp, medium confidence
Inkling-Small52.9%17.8%estimated ± 1.6 pp, medium confidence
K-Exaone32.1%8.0%estimated ± 1.6 pp, medium confidence
Kimi K2.661.8%24.2%estimated ± 1.6 pp, medium confidence
Kimi K2.546.8%14.3%estimated ± 1.6 pp, medium confidence
Kimi K2.5 (Reasoning)46.8%14.3%estimated ± 1.6 pp, medium confidence
Kimi K2.7 Code60.8%23.4%estimated ± 1.6 pp, medium confidence
Kimi K376.2%41.3%estimated ± 1.6 pp, high confidence
LFM2.5-2.6B7.7%1.5%estimated ± 1.6 pp, medium confidence
Ling 2.6 Flash25.3%5.8%estimated ± 1.6 pp, medium confidence
Ling 3.0 Flash50.6%16.4%estimated ± 1.6 pp, medium confidence
Ling 3.0 Flash FP850.6%16.4%estimated ± 1.6 pp, medium confidence
Llama 4 Maverick16.3%3.4%estimated ± 1.6 pp, medium confidence
Llama 4 Scout8.2%1.6%estimated ± 1.6 pp, medium confidence
MiMo-V2.5-Pro60.2%22.9%estimated ± 1.6 pp, medium confidence
MiMo-V2-Flash49.8%15.9%estimated ± 1.6 pp, medium confidence
MiniMax M2.752.6%17.6%estimated ± 1.6 pp, medium confidence
MiniMax M358.6%21.7%estimated ± 1.6 pp, medium confidence
Mistral Large 320.1%4.4%estimated ± 1.6 pp, medium confidence
Mistral Medium 3.5 128B46.9%14.3%estimated ± 1.6 pp, medium confidence
Mistral Small 426.6%6.3%estimated ± 1.6 pp, medium confidence
Mistral Small 4 (Reasoning)26.6%6.3%estimated ± 1.6 pp, medium confidence
Muse Glimmer 30B49.0%15.5%estimated ± 1.6 pp, medium confidence
Muse Spark58.6%21.7%estimated ± 1.6 pp, medium confidence
Muse Spark 1.171.3%34.2%estimated ± 1.6 pp, medium confidence
Muse Spark 1.272.2%35.4%estimated ± 1.6 pp, high confidence
Nemotron 3.5 Lightning 30B A3B NVFP426.8%6.3%estimated ± 1.6 pp, medium confidence
Nemotron 3 Nano 30B14.4%3.0%estimated ± 1.6 pp, medium confidence
Nemotron 3 Nano Omni 30B A3B13.8%2.8%estimated ± 1.6 pp, medium confidence
Nemotron 3 Super 100B37.7%10.1%estimated ± 1.6 pp, medium confidence
Nemotron 3 Ultra49.3%15.6%estimated ± 1.6 pp, medium confidence
o139.7%11.0%estimated ± 1.6 pp, medium confidence
o1-preview34.1%8.7%estimated ± 1.6 pp, medium confidence
Quasar 438B61.2%23.8%estimated ± 1.6 pp, medium confidence
Qwen3.5-122B-A10B45.7%13.7%estimated ± 1.6 pp, medium confidence
Qwen3.6-27B53.7%18.3%estimated ± 1.6 pp, medium confidence
Qwen3.6-35B-A3B41.9%11.9%estimated ± 1.6 pp, medium confidence
Qwen3.6 Plus54.5%18.8%estimated ± 1.6 pp, medium confidence
Qwen3.7 Max66.0%28.1%estimated ± 1.6 pp, medium confidence
Qwen3.7 Plus55.9%19.7%estimated ± 1.6 pp, medium confidence
Qwen3.8-27B68.1%30.4%estimated ± 1.6 pp, medium confidence
Qwen3.8-Flash-Next73.1%36.5%estimated ± 1.6 pp, high confidence
Qwen3.8 Max Preview71.8%34.8%estimated ± 1.6 pp, high confidence
Step 3.7 Flash39.6%10.9%estimated ± 1.6 pp, medium confidence
Trinity-Large-Preview25.8%6.0%estimated ± 1.6 pp, medium confidence
Trinity-Large-Thinking25.8%6.0%estimated ± 1.6 pp, medium confidence
Ultravox v0.6 Llama 3.3 70B11.9%2.4%estimated ± 1.6 pp, medium confidence