benchgap
Calibration

AA Coding Index → CursorBench 3.2

CursorBench 3.2 is estimated from AA Coding Index with a inverse Michaelis–Menten curve fitted on 17 models measured on both: y = 0.94752·(x − 0.0000) / (0.0000 + 1.8579 − x), R² = 0.72, cross-validated error 4.4 pp. It is used for 53 estimates.

Estimated modelAA Coding IndexCursorBench 3.2Source
Apodex 1.160.8%46.0%estimated ± 4.4 pp, high confidence
Apodex 1.1 Mini60.8%46.0%estimated ± 4.4 pp, high confidence
Celeris-114.4%8.0%estimated ± 4.4 pp, medium confidence
Claude 3 Opus19.5%11.1%estimated ± 4.4 pp, medium confidence
Command A+27.9%16.7%estimated ± 4.4 pp, medium confidence
DeepSeek V323.0%13.4%estimated ± 4.4 pp, medium confidence
Gemini 1.5 Pro23.6%13.8%estimated ± 4.4 pp, medium confidence
Gemini 2.5 Pro33.3%20.7%estimated ± 4.4 pp, medium confidence
Gemini 3.1 Pro68.8%55.8%estimated ± 4.4 pp, high confidence
Gemma 3 27B10.1%5.4%estimated ± 4.4 pp, medium confidence
Gemma 4 12B31.0%18.9%estimated ± 4.4 pp, medium confidence
Gemma 4 26B A4B39.3%25.4%estimated ± 4.4 pp, medium confidence
Gemma 4 31B43.4%28.9%estimated ± 4.4 pp, medium confidence
Gemma 4 E2B7.2%3.8%estimated ± 4.4 pp, medium confidence
Gemma 4 E4B9.4%5.0%estimated ± 4.4 pp, medium confidence
GLM-4.745.3%30.5%estimated ± 4.4 pp, medium confidence
GPT-4.1 mini20.2%11.6%estimated ± 4.4 pp, medium confidence
GPT-4.1 nano11.1%6.0%estimated ± 4.4 pp, medium confidence
GPT-4 Turbo21.5%12.4%estimated ± 4.4 pp, medium confidence
GPT-4o mini11.4%6.2%estimated ± 4.4 pp, medium confidence
GPT-5.149.4%34.3%estimated ± 4.4 pp, medium confidence
GPT-5.4 nano56.1%41.0%estimated ± 4.4 pp, medium confidence
GPT-5 (high)37.8%24.2%estimated ± 4.4 pp, medium confidence
GPT-OSS 120B30.4%18.6%estimated ± 4.4 pp, medium confidence
GPT-OSS 20B20.7%11.9%estimated ± 4.4 pp, medium confidence
Grok 4.342.3%27.9%estimated ± 4.4 pp, medium confidence
Hy358.8%43.9%estimated ± 4.4 pp, medium confidence
Hy3 Preview58.8%43.9%estimated ± 4.4 pp, medium confidence
K-Exaone32.1%19.8%estimated ± 4.4 pp, medium confidence
Kimi K2.5 (Reasoning)46.8%31.9%estimated ± 4.4 pp, medium confidence
LFM2.5-2.6B7.7%4.1%estimated ± 4.4 pp, medium confidence
Ling 2.6 Flash25.3%14.9%estimated ± 4.4 pp, medium confidence
Ling 3.0 Flash FP850.6%35.5%estimated ± 4.4 pp, medium confidence
Llama 4 Maverick16.3%9.1%estimated ± 4.4 pp, medium confidence
Llama 4 Scout8.2%4.4%estimated ± 4.4 pp, medium confidence
MiMo-V2-Flash49.8%34.7%estimated ± 4.4 pp, medium confidence
Mistral Large 320.1%11.5%estimated ± 4.4 pp, medium confidence
Mistral Medium 3.5 128B46.9%32.0%estimated ± 4.4 pp, medium confidence
Mistral Small 426.6%15.9%estimated ± 4.4 pp, medium confidence
Mistral Small 4 (Reasoning)26.6%15.9%estimated ± 4.4 pp, medium confidence
Nemotron 3.5 Lightning 30B A3B NVFP426.8%15.9%estimated ± 4.4 pp, medium confidence
Nemotron 3 Nano 30B14.4%7.9%estimated ± 4.4 pp, medium confidence
Nemotron 3 Nano Omni 30B A3B13.8%7.6%estimated ± 4.4 pp, medium confidence
Nemotron 3 Super 100B37.7%24.1%estimated ± 4.4 pp, medium confidence
Nemotron 3 Ultra49.3%34.2%estimated ± 4.4 pp, medium confidence
o139.7%25.8%estimated ± 4.4 pp, medium confidence
o1-preview34.1%21.3%estimated ± 4.4 pp, medium confidence
Quasar 438B61.2%46.5%estimated ± 4.4 pp, high confidence
Qwen3.5-122B-A10B45.7%30.9%estimated ± 4.4 pp, medium confidence
Qwen3.8 Max Preview71.8%59.7%estimated ± 4.4 pp, high confidence
Trinity-Large-Preview25.8%15.3%estimated ± 4.4 pp, medium confidence
Trinity-Large-Thinking25.8%15.3%estimated ± 4.4 pp, medium confidence
Ultravox v0.6 Llama 3.3 70B11.9%6.5%estimated ± 4.4 pp, medium confidence