benchgap
Calibration

CursorBench 3.2 → FrontierCode 1.1 Extended

FrontierCode 1.1 Extended is estimated from CursorBench 3.2 with a offset logistic curve fitted on 5 models measured on both: y = 0.5514 + (0.6246 − 0.5514) / (1 + exp(−152.87·(x − 0.6647))), R² = 0.95, cross-validated error 1.6 pp. It is used for 13 estimates.

Estimated modelCursorBench 3.2FrontierCode 1.1 ExtendedSource
Claude Fable 570.5%62.4%estimated ± 1.6 pp, medium confidence
Claude Fable 5.173.4%62.5%estimated ± 1.6 pp, low confidence
Claude Opus 4.862.3%55.2%estimated ± 1.6 pp, medium confidence
Claude Sonnet 561.5%55.1%estimated ± 1.6 pp, medium confidence
Composer 2.556.1%55.1%estimated ± 1.6 pp, low confidence
Gemini 3.5 Flash48.8%55.1%estimated ± 1.6 pp, low confidence
Gemini 3.6 Flash53.5%55.1%estimated ± 1.6 pp, low confidence
Gemini 3.8 Flash69.2%62.3%estimated ± 1.6 pp, medium confidence
GLM-5.255.0%55.1%estimated ± 1.6 pp, low confidence
GPT-5.558.4%55.1%estimated ± 1.6 pp, low confidence
Grok 4.566.7%59.4%estimated ± 1.6 pp, medium confidence
Kimi K2.7 Code49.7%55.1%estimated ± 1.6 pp, low confidence
Kimi K360.8%55.1%estimated ± 1.6 pp, low confidence