benchgap
Calibration

Vals SWE-bench → OpenHarmony Bench

OpenHarmony Bench is estimated from Vals SWE-bench with a linear curve fitted on 11 models measured on both: y = 0.3134·x + 0.2921, R² = 0.52, cross-validated error 3.2 pp. It is used for 13 estimates.

Estimated modelVals SWE-benchOpenHarmony BenchSource
Claude Haiku 4.566.6%50.1%estimated ± 3.2 pp, medium confidence
Claude Opus 4.782.0%54.9%estimated ± 3.2 pp, high confidence
Claude Sonnet 4.677.4%53.5%estimated ± 3.2 pp, high confidence
Composer 2.579.6%54.2%estimated ± 3.2 pp, high confidence
Gemini 3.1 Flash-Lite62.8%48.9%estimated ± 3.2 pp, medium confidence
Gemini 3 Flash75.0%52.7%estimated ± 3.2 pp, high confidence
GPT-5.2-Codex72.4%51.9%estimated ± 3.2 pp, high confidence
GPT-5.3 Codex78.0%53.7%estimated ± 3.2 pp, high confidence
Grok 4.2072.2%51.8%estimated ± 3.2 pp, high confidence
Laguna M.157.6%47.3%estimated ± 3.2 pp, medium confidence
Laguna XS.255.2%46.5%estimated ± 3.2 pp, medium confidence
MiMo-V2.571.0%51.5%estimated ± 3.2 pp, high confidence
Qwen3.5 Flash64.4%49.4%estimated ± 3.2 pp, medium confidence