Calibration
FrontierCode 1.1 Main → CursorBench 3.2
CursorBench 3.2 is estimated from FrontierCode 1.1 Main with a Michaelis–Menten curve fitted on 5 models measured on both: y = 2.0000·x / (1.00019 + x), R² = 0.93, cross-validated error 1.6 pp. It is used for 11 estimates.
| Estimated model | FrontierCode 1.1 Main | CursorBench 3.2 | Source |
|---|---|---|---|
| Claude Haiku 5.5 | 46.4% | 63.4% | estimated ± 1.6 pp, medium confidence |
| Claude Opus 4.6 | 26.9% | 42.4% | estimated ± 1.6 pp, low confidence |
| Claude Opus 4.7 | 38.5% | 55.6% | estimated ± 1.6 pp, low confidence |
| Claude Opus 5.5 | 54.4% | 70.5% | estimated ± 1.6 pp, low confidence |
| Claude Sonnet 4.6 | 24.3% | 39.1% | estimated ± 1.6 pp, low confidence |
| Claude Sonnet 5.5 | 46.2% | 63.2% | estimated ± 1.6 pp, medium confidence |
| Gemini 3.7 Flash | 43.6% | 60.7% | estimated ± 1.6 pp, medium confidence |
| GPT-5.4 mini | 27.0% | 42.5% | estimated ± 1.6 pp, low confidence |
| GPT-6 Astra | 53.3% | 69.5% | estimated ± 1.6 pp, medium confidence |
| SWE-1.7 | 42.3% | 59.4% | estimated ± 1.6 pp, low confidence |
| SWE-2 | 50.0% | 66.7% | estimated ± 1.6 pp, medium confidence |