benchgap
Coding

OpenHarmony Bench leaderboard

As of 2026-10-07, the highest measured score on OpenHarmony Bench is 60.8% by GLM-5.3. 157 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Mythos 560.9%estimated ± 4.9 pp, low confidence
2Sakana Fugu-Ultra60.9%estimated ± 4.9 pp, low confidence
3Beam60.9%estimated ± 4.9 pp, low confidence
4GLM-5.360.8%measured
5Qwen3.8 Max60.8%measured
6DeepSeek V4.1 Flash60.3%measured
7Ornith-1.5-397B59.3%estimated ± 2.7 pp, high confidence
8Hy4 preview59.1%estimated ± 2.7 pp, high confidence
9DeepSeek V4 Pro 081359.0%measured
10Claude Fable 558.8%estimated ± 2.4 pp, medium confidence
11Claude Fable 5.158.8%estimated ± 2.4 pp, medium confidence
12Claude Opus 558.8%estimated ± 2.4 pp, high confidence
13Claude Opus 5.558.8%estimated ± 2.4 pp, medium confidence
14Claude Sonnet 5.558.8%estimated ± 2.4 pp, medium confidence
15Gemini 3.1 Pro58.8%estimated ± 2.4 pp, high confidence
16Gemini 3.7 Flash58.8%estimated ± 2.4 pp, high confidence
17Gemini 3.8 Flash58.8%estimated ± 2.4 pp, high confidence
18Gemini 4 Argon58.8%estimated ± 2.4 pp, medium confidence
19GPT-5.6 Sol58.8%estimated ± 2.4 pp, high confidence
20GPT-6 Astra58.8%estimated ± 2.4 pp, high confidence
21GPT-6 Sol58.8%estimated ± 2.4 pp, high confidence
22Grok 4.658.8%estimated ± 2.4 pp, high confidence
23Grok 4.758.8%estimated ± 2.4 pp, high confidence
24MiMo-V2.6-Pro58.8%estimated ± 2.4 pp, medium confidence
25Muse Spark 1.158.8%estimated ± 2.4 pp, high confidence
26Muse Spark 1.258.8%estimated ± 2.4 pp, high confidence
27Muse Spark 1.358.8%estimated ± 2.4 pp, high confidence
28Step 5 Preview58.8%estimated ± 2.4 pp, high confidence
29GPT-5.558.8%estimated ± 2.4 pp, high confidence
30Claude Haiku 5.558.8%estimated ± 2.4 pp, high confidence
31GPT-5.6 Terra58.8%estimated ± 2.4 pp, high confidence
32Grok 4.558.8%estimated ± 2.4 pp, high confidence
33GPT-6 Luna58.8%estimated ± 2.4 pp, high confidence
34Claude Opus 4.858.8%estimated ± 2.4 pp, high confidence
35Claude Sonnet 558.8%estimated ± 2.4 pp, high confidence
36GPT-6.1 Sol58.8%estimated ± 2.4 pp, high confidence
37Mistral Large 458.8%estimated ± 2.4 pp, high confidence
38Ling 3.1 Flash58.8%estimated ± 2.4 pp, high confidence
39Gemini 3.5 Flash58.8%estimated ± 2.4 pp, high confidence
40GPT-5.6 Luna58.7%estimated ± 2.4 pp, high confidence
41Gemini 3.6 Flash58.7%estimated ± 2.4 pp, high confidence
42GPT-5.4 mini58.6%estimated ± 2.4 pp, high confidence
43GLM-5.258.4%measured
44Kimi K2.658.2%estimated ± 2.4 pp, high confidence
45Claude Opus 4.7 (Adaptive)58.1%estimated ± 2.8 pp, high confidence
46MiMo-V2.6-Flash58.0%estimated ± 2.4 pp, high confidence
47GLM-5.3-Flash57.3%measured
48Kimi K357.3%measured
49GPT-5.456.9%estimated ± 2.8 pp, high confidence
50MiMo-V2.5-Pro56.4%estimated ± 2.4 pp, high confidence
51Qwen3.8-Flash-Next56.4%estimated ± 2.4 pp, high confidence
52Qwen3.8 Max Preview56.0%measured
53dots3-note Preview55.2%estimated ± 2.7 pp, high confidence
54Claude Opus 4.754.9%estimated ± 3.2 pp, high confidence
55Qwen3.8-Omni-Flash54.7%estimated ± 2.7 pp, high confidence
56MiniMax M2.754.5%estimated ± 2.4 pp, high confidence
57Ornith-1.0-397B54.4%estimated ± 2.7 pp, high confidence
58Composer 2.554.2%estimated ± 3.2 pp, high confidence
59DeepSeek V4 Flash 073153.8%measured
60Seed 2.1 Pro53.8%estimated ± 2.7 pp, high confidence
61GPT-5.3 Codex53.7%estimated ± 3.2 pp, high confidence
62Claude Sonnet 4.653.5%estimated ± 3.2 pp, high confidence
63Qwen3.7 Max53.4%measured
64Ornith-1.5-35B-A3B53.4%estimated ± 2.7 pp, high confidence
65Atria Dawn Preview53.4%estimated ± 4.9 pp, low confidence
66Laguna S 2.153.3%estimated ± 4.9 pp, low confidence
67Sakana Fugu53.3%estimated ± 4.9 pp, low confidence
68Claude Opus 4.653.3%estimated ± 4.9 pp, low confidence
69GLM-553.3%estimated ± 4.9 pp, low confidence
70GPT-5.253.3%estimated ± 4.9 pp, low confidence
71Laguna XS 2.153.3%estimated ± 4.9 pp, low confidence
72LLaDA2.2-flash53.3%estimated ± 4.9 pp, low confidence
73LongCat-Flash-Lite-Sparse53.3%estimated ± 4.9 pp, low confidence
74MAI-Thinking-153.3%estimated ± 4.9 pp, low confidence
75Qwen3.5 397B53.3%estimated ± 4.9 pp, low confidence
76Inkling-Small53.1%estimated ± 2.4 pp, high confidence
77Gemini 3 Flash52.7%estimated ± 3.2 pp, high confidence
78GLM-5.152.3%measured
79Kimi K2.7 Code52.1%measured
80Seed 2.1 Turbo52.0%estimated ± 2.7 pp, high confidence
81GPT-5.2-Codex51.9%estimated ± 3.2 pp, high confidence
82Grok 4.2051.8%estimated ± 3.2 pp, high confidence
83Claude Opus 4.551.8%estimated ± 2.7 pp, high confidence
84Qwen 3.6 Max (preview)51.6%estimated ± 2.7 pp, high confidence
85MiMo-V2.551.5%estimated ± 3.2 pp, high confidence
86Muse Spark51.4%estimated ± 2.8 pp, high confidence
87Hy351.3%estimated ± 2.4 pp, high confidence
88Hy3 Preview51.3%estimated ± 2.4 pp, high confidence
89Grok 4.351.2%estimated ± 2.4 pp, high confidence
90Quasar 438B51.1%estimated ± 2.4 pp, high confidence
91GPT-5.4 nano51.0%estimated ± 2.4 pp, high confidence
92Inkling51.0%estimated ± 2.4 pp, high confidence
93Qwen3.8-27B51.0%estimated ± 2.4 pp, high confidence
94Gemini 2.5 Pro51.0%estimated ± 2.4 pp, high confidence
95Qwen3.7 Plus51.0%estimated ± 2.4 pp, high confidence
96Apodex 1.151.0%estimated ± 2.4 pp, high confidence
97Apodex 1.1 Mini51.0%estimated ± 2.4 pp, high confidence
98Gemma 4 31B51.0%estimated ± 2.4 pp, high confidence
99Muse Glimmer 30B51.0%estimated ± 2.4 pp, high confidence
100Solar Pro 451.0%estimated ± 2.4 pp, medium confidence
101A.X K251.0%estimated ± 2.4 pp, medium confidence
102Celeris-151.0%estimated ± 2.4 pp, medium confidence
103Command A+51.0%estimated ± 2.4 pp, medium confidence
104DeepSeek V351.0%estimated ± 2.4 pp, medium confidence
105DeepSeek V3 032451.0%estimated ± 2.4 pp, medium confidence
106Gemini 3.5 Flash-Lite51.0%estimated ± 2.4 pp, medium confidence
107Gemma 3 27B51.0%estimated ± 2.4 pp, medium confidence
108Gemma 4 26B A4B51.0%estimated ± 2.4 pp, medium confidence
109Gemma 4 E4B51.0%estimated ± 2.4 pp, medium confidence
110GPT-OSS 120B51.0%estimated ± 2.4 pp, medium confidence
111GPT-OSS 20B51.0%estimated ± 2.4 pp, medium confidence
112Granite 4.2 30B51.0%estimated ± 2.4 pp, medium confidence
113Granite 4.2 3B51.0%estimated ± 2.4 pp, medium confidence
114Granite 4.2 8B51.0%estimated ± 2.4 pp, medium confidence
115K-EXAONE 2.051.0%estimated ± 2.4 pp, medium confidence
116LFM2.5-2.6B51.0%estimated ± 2.4 pp, medium confidence
117Ling 3.0 Flash51.0%estimated ± 2.4 pp, medium confidence
118Ling 3.0 Flash FP851.0%estimated ± 2.4 pp, medium confidence
119Ling 3.0 Flash VL51.0%estimated ± 2.4 pp, medium confidence
120Ling 3.0 Tiny51.0%estimated ± 2.4 pp, medium confidence
121Llama 4 Maverick51.0%estimated ± 2.4 pp, medium confidence
122Llama 4 Scout51.0%estimated ± 2.4 pp, medium confidence
123Mercury 2.551.0%estimated ± 2.4 pp, medium confidence
124MiniCPM5-2B51.0%estimated ± 2.4 pp, medium confidence
125Mistral Large 351.0%estimated ± 2.4 pp, medium confidence
126Mistral Medium 3.5 128B51.0%estimated ± 2.4 pp, medium confidence
127Mistral Small 451.0%estimated ± 2.4 pp, medium confidence
128Mistral Small 4 (Reasoning)51.0%estimated ± 2.4 pp, medium confidence
129Nemotron 3.5 Lightning 30B A3B NVFP451.0%estimated ± 2.4 pp, medium confidence
130Nemotron 3 Nano 30B51.0%estimated ± 2.4 pp, medium confidence
131Nemotron 3 Super 100B51.0%estimated ± 2.4 pp, medium confidence
132Nemotron 3 Ultra51.0%estimated ± 2.4 pp, medium confidence
133North Mini Code51.0%estimated ± 2.4 pp, medium confidence
134Qwen3.5-122B-A10B51.0%estimated ± 2.4 pp, medium confidence
135Qwen3.6-27B51.0%estimated ± 2.4 pp, medium confidence
136Qwen3.6-35B-A3B51.0%estimated ± 2.4 pp, medium confidence
137Solar Pro 351.0%estimated ± 2.4 pp, medium confidence
138Step 3.7 Flash51.0%estimated ± 2.4 pp, medium confidence
139Trinity-Large-Preview51.0%estimated ± 2.4 pp, medium confidence
140Trinity-Large-Thinking51.0%estimated ± 2.4 pp, medium confidence
141Claude Haiku 4.550.1%estimated ± 3.2 pp, medium confidence
142Qwen3.6 Plus49.5%estimated ± 2.8 pp, medium confidence
143Qwen3.5 Flash49.4%estimated ± 3.2 pp, medium confidence
144Gemini 3.1 Flash-Lite48.9%estimated ± 3.2 pp, medium confidence
145MiniMax M348.4%measured
146MiMo-V2-Flash47.4%estimated ± 2.8 pp, medium confidence
147Laguna M.147.3%estimated ± 3.2 pp, medium confidence
148GPT-5.147.2%estimated ± 2.8 pp, medium confidence
149Laguna XS.246.5%estimated ± 3.2 pp, medium confidence
150Ornith-1.0-35B46.4%estimated ± 2.7 pp, medium confidence
151Kimi K2.546.0%estimated ± 2.8 pp, medium confidence
152Kimi K2.5 (Reasoning)46.0%estimated ± 2.8 pp, medium confidence
153GLM-4.745.3%estimated ± 2.8 pp, medium confidence
154Ornith-1.5-9B44.8%estimated ± 2.7 pp, medium confidence
155o142.8%estimated ± 2.8 pp, medium confidence
156GPT-5 (high)42.0%estimated ± 2.8 pp, medium confidence
157Ornith-1.0-9B40.7%estimated ± 2.7 pp, medium confidence
158o1-preview40.3%estimated ± 2.8 pp, medium confidence
159K-Exaone39.4%estimated ± 2.8 pp, medium confidence
160Gemma 4 12B38.9%estimated ± 2.8 pp, medium confidence
161Ling 2.6 Flash36.3%estimated ± 2.8 pp, medium confidence
162Gemini 1.5 Pro35.6%estimated ± 2.8 pp, medium confidence
163GPT-4 Turbo34.6%estimated ± 2.8 pp, medium confidence
164GPT-4.1 mini34.1%estimated ± 2.8 pp, medium confidence
165Claude 3 Opus33.7%estimated ± 2.8 pp, medium confidence
166Nemotron 3 Nano Omni 30B A3B31.1%estimated ± 2.8 pp, medium confidence
167Ultravox v0.6 Llama 3.3 70B30.3%estimated ± 2.8 pp, medium confidence
168GPT-4o mini30.1%estimated ± 2.8 pp, medium confidence
169GPT-4.1 nano30.0%estimated ± 2.8 pp, medium confidence
170Gemma 4 E2B28.2%estimated ± 2.8 pp, medium confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General