Calibration
CursorBench 3.2 → FrontierCode 1.1 Extended
FrontierCode 1.1 Extended is estimated from CursorBench 3.2 with a offset logistic curve fitted on 5 models measured on both: y = 0.5514 + (0.6246 − 0.5514) / (1 + exp(−152.87·(x − 0.6647))), R² = 0.95, cross-validated error 1.6 pp. It is used for 13 estimates.
| Estimated model | CursorBench 3.2 | FrontierCode 1.1 Extended | Source |
|---|---|---|---|
| Claude Fable 5 | 70.5% | 62.4% | estimated ± 1.6 pp, medium confidence |
| Claude Fable 5.1 | 73.4% | 62.5% | estimated ± 1.6 pp, low confidence |
| Claude Opus 4.8 | 62.3% | 55.2% | estimated ± 1.6 pp, medium confidence |
| Claude Sonnet 5 | 61.5% | 55.1% | estimated ± 1.6 pp, medium confidence |
| Composer 2.5 | 56.1% | 55.1% | estimated ± 1.6 pp, low confidence |
| Gemini 3.5 Flash | 48.8% | 55.1% | estimated ± 1.6 pp, low confidence |
| Gemini 3.6 Flash | 53.5% | 55.1% | estimated ± 1.6 pp, low confidence |
| Gemini 3.8 Flash | 69.2% | 62.3% | estimated ± 1.6 pp, medium confidence |
| GLM-5.2 | 55.0% | 55.1% | estimated ± 1.6 pp, low confidence |
| GPT-5.5 | 58.4% | 55.1% | estimated ± 1.6 pp, low confidence |
| Grok 4.5 | 66.7% | 59.4% | estimated ± 1.6 pp, medium confidence |
| Kimi K2.7 Code | 49.7% | 55.1% | estimated ± 1.6 pp, low confidence |
| Kimi K3 | 60.8% | 55.1% | estimated ± 1.6 pp, low confidence |