Calibration
CursorBench 3.2 → FrontierCode 1.1 Main
FrontierCode 1.1 Main is estimated from CursorBench 3.2 with a inverse Michaelis–Menten curve fitted on 5 models measured on both: y = 0.99814·(x − 0.0000) / (0.0000 + 2.0000 − x), R² = 0.92, cross-validated error 1.6 pp. It is used for 13 estimates.
| Estimated model | CursorBench 3.2 | FrontierCode 1.1 Main | Source |
|---|---|---|---|
| Claude Fable 5.1 | 73.4% | 57.9% | estimated ± 1.6 pp, low confidence |
| Composer 2.5 | 56.1% | 38.9% | estimated ± 1.6 pp, low confidence |
| Gemini 3.5 Flash | 48.8% | 32.2% | estimated ± 1.6 pp, low confidence |
| Gemini 3.6 Flash | 53.5% | 36.5% | estimated ± 1.6 pp, low confidence |
| Gemini 3.8 Flash | 69.2% | 52.8% | estimated ± 1.6 pp, medium confidence |
| GLM-5.2 | 55.0% | 37.9% | estimated ± 1.6 pp, low confidence |
| GPT-5.6 Luna | 61.1% | 43.9% | estimated ± 1.6 pp, medium confidence |
| GPT-5.6 Sol | 67.2% | 50.5% | estimated ± 1.6 pp, medium confidence |
| GPT-5.6 Terra | 64.9% | 47.9% | estimated ± 1.6 pp, medium confidence |
| Grok 4.5 | 66.7% | 49.9% | estimated ± 1.6 pp, medium confidence |
| Grok 4.6 | 70.8% | 54.7% | estimated ± 1.6 pp, low confidence |
| Kimi K2.7 Code | 49.7% | 33.0% | estimated ± 1.6 pp, low confidence |
| Kimi K3 | 60.8% | 43.6% | estimated ± 1.6 pp, medium confidence |