Calibration
DeepSWE → FrontierCode 1.1 Main
FrontierCode 1.1 Main is estimated from DeepSWE with a inverse Michaelis–Menten curve fitted on 6 models measured on both: y = 0.90750·(x − 0.0000) / (0.0000 + 2.0000 − x), R² = 0.45, cross-validated error 3.7 pp. It is used for 21 estimates.
| Estimated model | DeepSWE | FrontierCode 1.1 Main | Source |
|---|---|---|---|
| DeepSeek V4.1 Flash | 74.2% | 53.5% | estimated ± 3.7 pp, low confidence |
| GLM-5.3 | 66.9% | 45.6% | estimated ± 3.7 pp, low confidence |
| GLM-5.3-Flash | 63.4% | 42.1% | estimated ± 3.7 pp, low confidence |
| GPT-6.1 Sol | 71.9% | 50.9% | estimated ± 3.7 pp, low confidence |
| GPT-6 Luna | 66.6% | 45.3% | estimated ± 3.7 pp, low confidence |
| GPT-6 Sol | 68.8% | 47.6% | estimated ± 3.7 pp, low confidence |
| Grok 4.7 | 71.0% | 49.9% | estimated ± 3.7 pp, low confidence |
| Hy4 preview | 64.3% | 43.0% | estimated ± 3.7 pp, low confidence |
| Laguna S 2.1 | 40.4% | 23.0% | estimated ± 3.7 pp, low confidence |
| MiMo-V2.6-Flash | 67.9% | 46.6% | estimated ± 3.7 pp, low confidence |
| MiMo-V2.6-Pro | 71.9% | 50.9% | estimated ± 3.7 pp, low confidence |
| Muse Spark 1.1 | 53.3% | 33.0% | estimated ± 3.7 pp, low confidence |
| Muse Spark 1.2 | 59.3% | 38.2% | estimated ± 3.7 pp, low confidence |
| Muse Spark 1.3 | 75.4% | 54.9% | estimated ± 3.7 pp, low confidence |
| Pareto 26.10 Preview | 69.9% | 48.8% | estimated ± 3.7 pp, low confidence |
| Pareto 26.9 | 74.0% | 53.3% | estimated ± 3.7 pp, low confidence |
| Qwen3.8-27B | 42.2% | 24.3% | estimated ± 3.7 pp, low confidence |
| Qwen3.8-Flash-Next | 58.7% | 37.7% | estimated ± 3.7 pp, low confidence |
| Qwen3.8 Max | 56.6% | 35.8% | estimated ± 3.7 pp, low confidence |
| Qwen3.8-Omni-Flash | 57.8% | 36.9% | estimated ± 3.7 pp, low confidence |
| Step 5 Preview | 67.7% | 46.4% | estimated ± 3.7 pp, low confidence |