benchgap
Calibration

FrontierCode 1.1 Main → CursorBench 3.2

CursorBench 3.2 is estimated from FrontierCode 1.1 Main with a Michaelis–Menten curve fitted on 5 models measured on both: y = 2.0000·x / (1.00019 + x), R² = 0.93, cross-validated error 1.6 pp. It is used for 11 estimates.

Estimated modelFrontierCode 1.1 MainCursorBench 3.2Source
Claude Haiku 5.546.4%63.4%estimated ± 1.6 pp, medium confidence
Claude Opus 4.626.9%42.4%estimated ± 1.6 pp, low confidence
Claude Opus 4.738.5%55.6%estimated ± 1.6 pp, low confidence
Claude Opus 5.554.4%70.5%estimated ± 1.6 pp, low confidence
Claude Sonnet 4.624.3%39.1%estimated ± 1.6 pp, low confidence
Claude Sonnet 5.546.2%63.2%estimated ± 1.6 pp, medium confidence
Gemini 3.7 Flash43.6%60.7%estimated ± 1.6 pp, medium confidence
GPT-5.4 mini27.0%42.5%estimated ± 1.6 pp, low confidence
GPT-6 Astra53.3%69.5%estimated ± 1.6 pp, medium confidence
SWE-1.742.3%59.4%estimated ± 1.6 pp, low confidence
SWE-250.0%66.7%estimated ± 1.6 pp, medium confidence