benchgap
Calibration

AA Coding Index → CursorBench 3.1

CursorBench 3.1 is estimated from AA Coding Index with a offset logistic curve fitted on 5 models measured on both: y = 0.4791 + (1.2000 − 0.4791) / (1 + exp(−47.94·(x − 0.7815))), R² = 0.99, cross-validated error 4.8 pp. It is used for 24 estimates.

Estimated modelAA Coding IndexCursorBench 3.1Source
Claude 3 Opus19.5%47.9%estimated ± 4.8 pp, low confidence
Claude Opus 4.7 (Adaptive)73.6%55.2%estimated ± 4.8 pp, medium confidence
Gemini 1.5 Pro23.6%47.9%estimated ± 4.8 pp, low confidence
Gemma 4 12B31.0%47.9%estimated ± 4.8 pp, low confidence
Gemma 4 E2B7.2%47.9%estimated ± 4.8 pp, low confidence
GLM-4.745.3%47.9%estimated ± 4.8 pp, low confidence
GPT-4.1 mini20.2%47.9%estimated ± 4.8 pp, low confidence
GPT-4.1 nano11.1%47.9%estimated ± 4.8 pp, low confidence
GPT-4 Turbo21.5%47.9%estimated ± 4.8 pp, low confidence
GPT-4o mini11.4%47.9%estimated ± 4.8 pp, low confidence
GPT-5.149.4%47.9%estimated ± 4.8 pp, low confidence
GPT-5.471.1%50.2%estimated ± 4.8 pp, medium confidence
GPT-5 (high)37.8%47.9%estimated ± 4.8 pp, low confidence
K-Exaone32.1%47.9%estimated ± 4.8 pp, low confidence
Kimi K2.546.8%47.9%estimated ± 4.8 pp, low confidence
Kimi K2.5 (Reasoning)46.8%47.9%estimated ± 4.8 pp, low confidence
Ling 2.6 Flash25.3%47.9%estimated ± 4.8 pp, low confidence
MiMo-V2-Flash49.8%47.9%estimated ± 4.8 pp, low confidence
Muse Spark58.6%47.9%estimated ± 4.8 pp, low confidence
Nemotron 3 Nano Omni 30B A3B13.8%47.9%estimated ± 4.8 pp, low confidence
o139.7%47.9%estimated ± 4.8 pp, low confidence
o1-preview34.1%47.9%estimated ± 4.8 pp, low confidence
Qwen3.6 Plus54.5%47.9%estimated ± 4.8 pp, low confidence
Ultravox v0.6 Llama 3.3 70B11.9%47.9%estimated ± 4.8 pp, low confidence