benchgap
Calibration

SWE-bench Pro → OpenHarmony Bench

OpenHarmony Bench is estimated from SWE-bench Pro with a offset logistic curve fitted on 7 models measured on both: y = 0.5328 + (0.6090 − 0.5328) / (1 + exp(−200.00·(x − 0.6181))), R² = 0.50, cross-validated error 4.9 pp. It is used for 14 estimates.

Estimated modelSWE-bench ProOpenHarmony BenchSource
Atria Dawn Preview59.6%53.4%estimated ± 4.9 pp, low confidence
Claude Mythos 580.3%60.9%estimated ± 4.9 pp, low confidence
Claude Opus 4.653.4%53.3%estimated ± 4.9 pp, low confidence
GLM-555.1%53.3%estimated ± 4.9 pp, low confidence
GPT-5.255.6%53.3%estimated ± 4.9 pp, low confidence
Laguna S 2.159.4%53.3%estimated ± 4.9 pp, low confidence
Laguna XS 2.147.6%53.3%estimated ± 4.9 pp, low confidence
LLaDA2.2-flash30.1%53.3%estimated ± 4.9 pp, low confidence
LongCat-Flash-Lite-Sparse40.6%53.3%estimated ± 4.9 pp, low confidence
MAI-Thinking-152.8%53.3%estimated ± 4.9 pp, low confidence
Qwen3.5 397B50.9%53.3%estimated ± 4.9 pp, low confidence
Beam65.5%60.9%estimated ± 4.9 pp, low confidence
Sakana Fugu59.0%53.3%estimated ± 4.9 pp, low confidence
Sakana Fugu-Ultra73.7%60.9%estimated ± 4.9 pp, low confidence