benchgap
Calibration

AA Coding Index → OpenHarmony Bench

OpenHarmony Bench is estimated from AA Coding Index with a linear curve fitted on 10 models measured on both: y = 0.4503·x + 0.2495, R² = 0.65, cross-validated error 2.8 pp. It is used for 24 estimates.

Estimated modelAA Coding IndexOpenHarmony BenchSource
Claude 3 Opus19.5%33.7%estimated ± 2.8 pp, medium confidence
Claude Opus 4.7 (Adaptive)73.6%58.1%estimated ± 2.8 pp, high confidence
Gemini 1.5 Pro23.6%35.6%estimated ± 2.8 pp, medium confidence
Gemma 4 12B31.0%38.9%estimated ± 2.8 pp, medium confidence
Gemma 4 E2B7.2%28.2%estimated ± 2.8 pp, medium confidence
GLM-4.745.3%45.3%estimated ± 2.8 pp, medium confidence
GPT-4.1 mini20.2%34.1%estimated ± 2.8 pp, medium confidence
GPT-4.1 nano11.1%30.0%estimated ± 2.8 pp, medium confidence
GPT-4 Turbo21.5%34.6%estimated ± 2.8 pp, medium confidence
GPT-4o mini11.4%30.1%estimated ± 2.8 pp, medium confidence
GPT-5.149.4%47.2%estimated ± 2.8 pp, medium confidence
GPT-5.471.1%56.9%estimated ± 2.8 pp, high confidence
GPT-5 (high)37.8%42.0%estimated ± 2.8 pp, medium confidence
K-Exaone32.1%39.4%estimated ± 2.8 pp, medium confidence
Kimi K2.546.8%46.0%estimated ± 2.8 pp, medium confidence
Kimi K2.5 (Reasoning)46.8%46.0%estimated ± 2.8 pp, medium confidence
Ling 2.6 Flash25.3%36.3%estimated ± 2.8 pp, medium confidence
MiMo-V2-Flash49.8%47.4%estimated ± 2.8 pp, medium confidence
Muse Spark58.6%51.4%estimated ± 2.8 pp, high confidence
Nemotron 3 Nano Omni 30B A3B13.8%31.1%estimated ± 2.8 pp, medium confidence
o139.7%42.8%estimated ± 2.8 pp, medium confidence
o1-preview34.1%40.3%estimated ± 2.8 pp, medium confidence
Qwen3.6 Plus54.5%49.5%estimated ± 2.8 pp, medium confidence
Ultravox v0.6 Llama 3.3 70B11.9%30.3%estimated ± 2.8 pp, medium confidence