benchgap
Calibration

DeepSWE → FrontierCode 1.1 Main

FrontierCode 1.1 Main is estimated from DeepSWE with a inverse Michaelis–Menten curve fitted on 6 models measured on both: y = 0.90750·(x − 0.0000) / (0.0000 + 2.0000 − x), R² = 0.45, cross-validated error 3.7 pp. It is used for 21 estimates.

Estimated modelDeepSWEFrontierCode 1.1 MainSource
DeepSeek V4.1 Flash74.2%53.5%estimated ± 3.7 pp, low confidence
GLM-5.366.9%45.6%estimated ± 3.7 pp, low confidence
GLM-5.3-Flash63.4%42.1%estimated ± 3.7 pp, low confidence
GPT-6.1 Sol71.9%50.9%estimated ± 3.7 pp, low confidence
GPT-6 Luna66.6%45.3%estimated ± 3.7 pp, low confidence
GPT-6 Sol68.8%47.6%estimated ± 3.7 pp, low confidence
Grok 4.771.0%49.9%estimated ± 3.7 pp, low confidence
Hy4 preview64.3%43.0%estimated ± 3.7 pp, low confidence
Laguna S 2.140.4%23.0%estimated ± 3.7 pp, low confidence
MiMo-V2.6-Flash67.9%46.6%estimated ± 3.7 pp, low confidence
MiMo-V2.6-Pro71.9%50.9%estimated ± 3.7 pp, low confidence
Muse Spark 1.153.3%33.0%estimated ± 3.7 pp, low confidence
Muse Spark 1.259.3%38.2%estimated ± 3.7 pp, low confidence
Muse Spark 1.375.4%54.9%estimated ± 3.7 pp, low confidence
Pareto 26.10 Preview69.9%48.8%estimated ± 3.7 pp, low confidence
Pareto 26.974.0%53.3%estimated ± 3.7 pp, low confidence
Qwen3.8-27B42.2%24.3%estimated ± 3.7 pp, low confidence
Qwen3.8-Flash-Next58.7%37.7%estimated ± 3.7 pp, low confidence
Qwen3.8 Max56.6%35.8%estimated ± 3.7 pp, low confidence
Qwen3.8-Omni-Flash57.8%36.9%estimated ± 3.7 pp, low confidence
Step 5 Preview67.7%46.4%estimated ± 3.7 pp, low confidence