benchgap
Agentic · tools

Gert Labs leaderboard

As of 2026-10-07, the highest measured score on Gert Labs is 73.0% by Claude Opus 4.8. 133 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Fable 5.1100.0%estimated ± 3.0 pp, low confidence
2GPT-6 Astra100.0%estimated ± 3.0 pp, low confidence
3Claude Sonnet 5.585.6%estimated ± 8.9 pp, low confidence
4Qwen3.8 Max Preview81.6%estimated ± 8.5 pp, low confidence
5Claude Opus 581.4%estimated ± 3.0 pp, low confidence
6Grok 4.781.3%estimated ± 8.9 pp, low confidence
7Claude Fable 580.2%estimated ± 3.0 pp, low confidence
8Qwen3.8-27B78.5%estimated ± 7.2 pp, low confidence
9Gemini 4 Argon76.7%estimated ± 3.7 pp, low confidence
10Gemini 3.6 Flash76.4%estimated ± 7.2 pp, medium confidence
11GPT-6 Sol76.4%estimated ± 3.7 pp, low confidence
12Muse Spark 1.275.9%estimated ± 8.5 pp, low confidence
13GPT-6.1 Sol75.9%estimated ± 8.9 pp, low confidence
14GPT-5.5 Pro75.9%estimated ± 7.4 pp, low confidence
15Claude Opus 5.575.8%estimated ± 3.7 pp, low confidence
16Holo3-35B-A3B75.7%estimated ± 7.2 pp, medium confidence
17GPT-5.6 Sol75.5%estimated ± 3.0 pp, low confidence
18GPT-5.4 Pro74.7%estimated ± 7.4 pp, low confidence
19Gemini 3.8 Flash74.4%estimated ± 3.0 pp, low confidence
20Grok 4.574.0%estimated ± 8.5 pp, medium confidence
21MiMo-V2.6-Flash73.4%estimated ± 6.0 pp, low confidence
22MiMo-V2.6-Pro73.0%estimated ± 6.0 pp, low confidence
23Claude Opus 4.873.0%measured
24GPT-5.572.9%measured
25Qwen3.8-Flash-Next71.9%estimated ± 3.7 pp, medium confidence
26Qwen3.8 Max71.9%estimated ± 3.7 pp, medium confidence
27Muse Spark 1.371.5%estimated ± 3.0 pp, medium confidence
28Claude Opus 4.7 (Adaptive)71.5%estimated ± 3.7 pp, medium confidence
29DeepSeek V4.1 Flash70.9%estimated ± 6.0 pp, low confidence
30Kimi K370.9%estimated ± 3.0 pp, medium confidence
31Ling 3.1 Flash70.8%estimated ± 6.0 pp, low confidence
32Fugu Cyber70.5%estimated ± 6.0 pp, low confidence
33Atria Dawn Preview70.3%estimated ± 6.0 pp, low confidence
34Gemini 3.8 Flash Cyber70.2%estimated ± 6.0 pp, low confidence
35Holo3-122B-A10B70.1%estimated ± 7.2 pp, medium confidence
36GPT-6 Luna70.0%estimated ± 8.9 pp, medium confidence
37Claude Sonnet 569.7%estimated ± 3.0 pp, medium confidence
38Gemini 3.7 Flash69.7%estimated ± 3.0 pp, medium confidence
39GPT-5.6 Terra69.7%estimated ± 3.0 pp, medium confidence
40Step 5 Preview69.7%estimated ± 6.0 pp, low confidence
41Muse Spark 1.169.7%estimated ± 3.7 pp, medium confidence
42GLM-5.369.6%estimated ± 6.0 pp, low confidence
43Claude Mythos 569.3%estimated ± 6.0 pp, low confidence
44Mistral Large 469.3%estimated ± 8.9 pp, medium confidence
45DeepSeek V4 Pro 081369.2%estimated ± 6.0 pp, low confidence
46Gemini 3.5 Flash Cyber69.1%estimated ± 6.0 pp, low confidence
47Claude Mythos Preview69.1%estimated ± 6.0 pp, low confidence
48Grok 4.668.0%estimated ± 3.0 pp, medium confidence
49UI-Mate-27B67.4%estimated ± 7.2 pp, medium confidence
50Hy4 preview67.4%estimated ± 6.0 pp, low confidence
51DeepSeek V4 Flash 073166.8%estimated ± 6.0 pp, low confidence
52Claude Opus 4.765.6%measured
53GPT-5.464.9%measured
54Quasar 438B64.6%estimated ± 8.5 pp, medium confidence
55GPT-5.6 Luna64.5%estimated ± 3.0 pp, medium confidence
56Qwen3.7 Max64.3%measured
57Claude Opus 4.564.2%measured
58Kimi K2.7 Code63.9%estimated ± 8.2 pp, low confidence
59Gemini 3.5 Flash-Lite63.3%estimated ± 7.2 pp, medium confidence
60Gemini 3 Pro63.2%measured
61Claude Sonnet 4.662.9%measured
62K-EXAONE 2.062.7%estimated ± 9.1 pp, low confidence
63MiMo-V2.5-Pro62.7%measured
64Ornith-1.0-397B62.4%estimated ± 9.1 pp, low confidence
65Claude Opus 4.661.9%measured
66Gemini 3.5 Flash61.9%measured
67GLM-5.160.1%measured
68Claude Opus 4.6 (Adaptive)59.5%estimated ± 11.3 pp, low confidence
69Inkling-Small58.5%estimated ± 7.4 pp, medium confidence
70Beam58.5%estimated ± 7.4 pp, medium confidence
71Apodex 1.158.3%estimated ± 8.9 pp, medium confidence
72Apodex 1.1 Mini58.3%estimated ± 8.9 pp, medium confidence
73Ornith-1.0-35B58.2%estimated ± 9.1 pp, low confidence
74Inkling58.1%estimated ± 7.4 pp, medium confidence
75Solar Open 257.5%estimated ± 8.2 pp, low confidence
76Hy357.5%estimated ± 8.5 pp, medium confidence
77GPT-5.3 Codex57.5%measured
78Gemini 3.1 Pro56.9%measured
79Kimi K2.656.8%measured
80Ling 3.0 Flash VL56.7%estimated ± 8.9 pp, medium confidence
81Gemini 3 Flash56.6%measured
82MiniMax M356.2%estimated ± 3.7 pp, medium confidence
83Agents-A156.2%estimated ± 7.4 pp, medium confidence
84Ornith-1.5-397B55.5%estimated ± 6.1 pp, low confidence
85GLM-5.255.3%estimated ± 7.2 pp, medium confidence
86Qwen3.6-27B54.8%measured
87Muse Spark54.8%estimated ± 6.0 pp, low confidence
88dots3-note Preview54.5%estimated ± 6.1 pp, low confidence
89Ornith-1.0-9B54.1%estimated ± 9.1 pp, low confidence
90UI-Mate-9B53.3%estimated ± 7.2 pp, medium confidence
91LLaDA2.2-flash53.2%estimated ± 8.2 pp, low confidence
92Muse Glimmer 30B52.9%estimated ± 7.2 pp, medium confidence
93Ling 3.0 Flash FP852.8%estimated ± 8.5 pp, medium confidence
94Ling 3.0 Flash51.8%estimated ± 6.1 pp, low confidence
95GPT-5.2-Codex51.8%measured
96Step 3.7 Flash51.6%measured
97GLM-551.0%measured
98Claude 4.1 Opus50.8%estimated ± 7.3 pp, medium confidence
99GPT-5.4 mini50.7%estimated ± 7.2 pp, medium confidence
100Qwen3.6 Plus50.6%measured
101Qwen3.5 Plus50.4%estimated ± 7.3 pp, medium confidence
102LLaDA2.2-mini50.3%estimated ± 9.1 pp, low confidence
103Claude Haiku 4.550.3%estimated ± 7.3 pp, medium confidence
104GPT-5 (high)50.2%estimated ± 7.3 pp, low confidence
105GPT-5.1-Codex49.7%measured
106GLM-5-Turbo49.4%estimated ± 9.1 pp, low confidence
107Grok Build 0.149.2%measured
108Ornith-1.5-35B-A3B48.7%estimated ± 6.1 pp, low confidence
109Claude Sonnet 4.548.5%measured
110Grok 4.148.3%estimated ± 8.5 pp, low confidence
111Qwen3.7 Plus47.5%estimated ± 3.7 pp, low confidence
112Grok 4.1 Fast47.3%measured
113GPT-5.4 nano47.2%estimated ± 7.2 pp, medium confidence
114A.X K247.2%estimated ± 8.9 pp, medium confidence
115MiMo-V2.546.9%measured
116Qwen3.5 397B46.8%measured
117GPT-5.246.5%measured
118Agents-A1-4B46.4%estimated ± 7.4 pp, medium confidence
119Kimi K2.545.9%measured
120Ornith-1.5-9B44.1%estimated ± 6.1 pp, low confidence
121Qwen3.5-122B-A10B44.0%estimated ± 7.2 pp, medium confidence
122Grok 4.343.9%measured
123Qwen3 Max43.7%measured
124Qwen3.6-35B-A3B42.7%measured
125Grok 442.3%measured
126Gemini 2.5 Pro42.0%measured
127MiMo-V2-Omni42.0%estimated ± 9.1 pp, low confidence
128GPT-5.141.2%measured
129MiniMax M2.740.4%measured
130GLM-4.740.0%measured
131Claude 4 Sonnet39.7%measured
132Qwen3.5-27B39.4%measured
133Mistral Medium 3.5 128B39.1%measured
134Gemini 3.1 Flash-Lite38.5%measured
135Grok 4.2038.4%measured
136MiniCPM5-2B38.3%estimated ± 8.9 pp, medium confidence
137Hy3 Preview36.9%measured
138MiMo-V2-Flash36.9%estimated ± 8.9 pp, medium confidence
139MiMo-V2-Pro36.7%measured
140Granite 4.2 30B36.3%estimated ± 8.9 pp, medium confidence
141Gemma 4 26B A4B36.2%estimated ± 8.9 pp, medium confidence
142Ling 3.0 Tiny36.2%estimated ± 8.9 pp, medium confidence
143DeepSeek V3 032436.0%estimated ± 8.9 pp, medium confidence
144Gemma 4 12B36.0%estimated ± 8.9 pp, medium confidence
145Gemma 4 E2B36.0%estimated ± 8.9 pp, medium confidence
146Gemma 4 E4B36.0%estimated ± 8.9 pp, medium confidence
147GPT-4.1 mini36.0%estimated ± 8.9 pp, medium confidence
148GPT-4.1 nano36.0%estimated ± 8.9 pp, medium confidence
149GPT-4o36.0%estimated ± 8.9 pp, medium confidence
150GPT-4o mini36.0%estimated ± 8.9 pp, medium confidence
151Granite 4.2 3B36.0%estimated ± 8.9 pp, medium confidence
152K-Exaone36.0%estimated ± 8.9 pp, medium confidence
153LFM2.5-2.6B36.0%estimated ± 8.9 pp, medium confidence
154Ling 2.6 Flash36.0%estimated ± 8.9 pp, medium confidence
155Mercury 2.536.0%estimated ± 8.9 pp, medium confidence
156Nemotron 3 Nano Omni 30B A3B36.0%estimated ± 8.9 pp, medium confidence
157North Mini Code36.0%estimated ± 8.9 pp, medium confidence
158Solar Pro 336.0%estimated ± 8.9 pp, medium confidence
159Ultravox v0.6 Llama 3.3 70B36.0%estimated ± 8.9 pp, medium confidence
160Nemotron 3 Super 100B35.8%estimated ± 8.5 pp, medium confidence
161Granite 4.2 8B35.4%estimated ± 8.5 pp, medium confidence
162Command A+35.3%estimated ± 8.5 pp, medium confidence
163Gemma 4 31B35.3%measured
164Mistral Large 334.1%estimated ± 8.5 pp, medium confidence
165Mistral Small 433.1%estimated ± 8.5 pp, medium confidence
166Mistral Small 4 (Reasoning)33.1%estimated ± 8.5 pp, medium confidence
167GPT-OSS 20B33.1%estimated ± 8.5 pp, medium confidence
168Trinity-Large-Preview32.9%estimated ± 8.5 pp, medium confidence
169Nemotron 3 Nano 30B32.7%estimated ± 8.5 pp, low confidence
170Kimi K2.5 (Reasoning)32.6%measured
171Trinity-Large-Thinking32.6%measured
172DeepSeek V332.5%estimated ± 8.5 pp, low confidence
173Celeris-132.3%estimated ± 8.5 pp, low confidence
174Llama 4 Maverick32.3%estimated ± 8.5 pp, low confidence
175Llama 4 Scout32.2%estimated ± 8.5 pp, low confidence
176Gemma 3 27B31.8%estimated ± 8.5 pp, low confidence
177GLM-5V-Turbo30.8%measured
178Solar Pro 430.2%estimated ± 7.4 pp, low confidence
179LongCat-Flash-Lite-Sparse29.7%estimated ± 7.4 pp, low confidence
180GPT-OSS 120B29.6%measured
181DeepSeek V3.229.6%measured
182Qwen3.5-35B-A3B29.0%measured
183Nemotron 3 Ultra26.4%estimated ± 7.4 pp, low confidence
184GPT-4.125.7%measured
185Nemotron 3.5 Lightning 30B A3B NVFP420.9%estimated ± 7.4 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General