benchgap
Agentic · tools

AA-AnalystAgent leaderboard

As of 2026-10-07, the highest measured score on AA-AnalystAgent is 60.0% by Gemini 3.7 Flash. 131 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Gemini 3.7 Flash60.0%measured
2Claude Fable 5.157.5%measured
3Claude Sonnet 5.557.5%measured
4Claude Opus 5.556.3%measured
5DeepSeek V4.1 Flash55.6%estimated ± 2.5 pp, medium confidence
6GLM-5.355.1%estimated ± 2.5 pp, medium confidence
7Gemini 4 Argon54.4%estimated ± 5.5 pp, medium confidence
8Muse Spark 1.354.4%estimated ± 5.5 pp, medium confidence
9GPT-6 Sol54.4%estimated ± 5.5 pp, medium confidence
10Gemini 3.8 Flash54.4%estimated ± 5.5 pp, medium confidence
11GPT-5.6 Terra54.4%estimated ± 5.5 pp, medium confidence
12GPT-5.6 Luna54.4%estimated ± 5.5 pp, medium confidence
13Claude Opus 4.854.1%estimated ± 5.5 pp, medium confidence
14Qwen3.8-Flash-Next53.9%estimated ± 5.5 pp, medium confidence
15DeepSeek V4 Pro 081353.9%estimated ± 2.5 pp, medium confidence
16Claude Opus 553.8%measured
17Claude Opus 4.7 (Adaptive)53.7%estimated ± 5.5 pp, medium confidence
18GLM-5.252.7%estimated ± 8.5 pp, low confidence
19Muse Spark 1.151.7%estimated ± 5.5 pp, medium confidence
20Claude Haiku 5.551.5%estimated ± 2.5 pp, medium confidence
21Claude Sonnet 551.5%estimated ± 2.5 pp, medium confidence
22Claude Opus 4.751.4%estimated ± 5.5 pp, medium confidence
23GPT-6 Astra51.2%measured
24Qwen3.8 Max Preview50.6%estimated ± 12.0 pp, low confidence
25GPT-6 Luna50.1%estimated ± 6.5 pp, medium confidence
26GPT-5.550.0%measured
27GPT-6.1 Sol50.0%measured
28Qwen3.8 Max49.8%estimated ± 2.5 pp, medium confidence
29Apodex 1.149.6%estimated ± 2.5 pp, medium confidence
30Ornith-1.5-397B49.6%estimated ± 2.5 pp, medium confidence
31Claude Fable 548.8%measured
32Hy4 preview48.4%estimated ± 2.5 pp, medium confidence
33GLM-5.3-Flash48.2%estimated ± 2.5 pp, medium confidence
34Ling 3.1 Flash47.6%estimated ± 12.0 pp, low confidence
35GPT-5.6 Sol47.5%measured
36Atria Dawn Preview47.4%estimated ± 12.5 pp, low confidence
37Grok 4.746.7%estimated ± 6.5 pp, medium confidence
38MiMo-V2.6-Pro45.3%estimated ± 6.5 pp, medium confidence
39Gemini 3.5 Flash45.2%estimated ± 3.6 pp, medium confidence
40Qwen3.7 Max44.3%estimated ± 2.5 pp, medium confidence
41Mistral Large 444.0%estimated ± 6.5 pp, medium confidence
42GPT-5.5 Pro43.2%estimated ± 12.5 pp, low confidence
43MiMo-V2.6-Flash42.5%estimated ± 6.7 pp, low confidence
44dots3-note Preview41.9%estimated ± 2.5 pp, medium confidence
45Grok 4.641.8%estimated ± 3.6 pp, medium confidence
46GPT-5.4 Pro41.8%estimated ± 12.5 pp, low confidence
47Claude Mythos 539.6%estimated ± 12.5 pp, low confidence
48Muse Spark 1.239.5%estimated ± 12.0 pp, low confidence
49Kimi K338.8%measured
50Step 5 Preview35.0%measured
51Grok 4.534.9%estimated ± 12.0 pp, low confidence
52Qwen3.8-27B34.2%estimated ± 3.6 pp, medium confidence
53Claude Opus 4.632.5%estimated ± 12.5 pp, low confidence
54Gemini 3.6 Flash29.9%estimated ± 12.0 pp, low confidence
55GPT-5.428.1%estimated ± 12.0 pp, low confidence
56Claude Sonnet 4.626.7%estimated ± 5.5 pp, medium confidence
57Hy3 Preview26.6%estimated ± 12.0 pp, low confidence
58Apodex 1.1 Mini25.7%estimated ± 12.0 pp, low confidence
59Ling 3.0 Flash25.6%estimated ± 9.1 pp, low confidence
60Quasar 438B25.2%estimated ± 12.0 pp, low confidence
61Agents-A125.2%estimated ± 2.5 pp, medium confidence
62Ling 3.0 Flash VL24.3%estimated ± 12.0 pp, low confidence
63Step 3.7 Flash23.8%estimated ± 2.5 pp, medium confidence
64Inkling23.8%measured
65Beam23.5%estimated ± 12.5 pp, low confidence
66MiMo-V2.5-Pro22.6%estimated ± 12.0 pp, low confidence
67GLM-5.122.4%estimated ± 12.0 pp, low confidence
68Solar Pro 422.0%estimated ± 12.0 pp, low confidence
69Grok 4.320.9%estimated ± 12.0 pp, low confidence
70Hy320.1%estimated ± 12.0 pp, low confidence
71Kimi K2.7 Code19.1%estimated ± 12.0 pp, low confidence
72GLM-4.718.1%estimated ± 12.0 pp, low confidence
73GPT-5.4 mini18.1%estimated ± 12.0 pp, low confidence
74MiniMax M2.718.0%estimated ± 12.0 pp, low confidence
75Muse Spark17.5%estimated ± 12.0 pp, low confidence
76DeepSeek V4 Flash 073117.5%estimated ± 2.5 pp, medium confidence
77Qwen3.6 Plus17.2%estimated ± 12.0 pp, low confidence
78Qwen3.6-27B17.0%estimated ± 12.0 pp, low confidence
79Gemini 3.5 Flash-Lite16.9%estimated ± 12.0 pp, low confidence
80Muse Glimmer 30B16.4%estimated ± 3.6 pp, medium confidence
81A.X K216.0%estimated ± 12.0 pp, low confidence
82GPT-5.4 nano15.5%estimated ± 12.0 pp, low confidence
83Ling 3.0 Flash FP815.4%estimated ± 12.0 pp, low confidence
84GPT-5 (high)14.3%estimated ± 12.0 pp, low confidence
85Qwen3.6-35B-A3B13.4%estimated ± 12.0 pp, low confidence
86Agents-A1-4B12.7%estimated ± 12.5 pp, low confidence
87GPT-5.212.0%estimated ± 12.5 pp, low confidence
88Kimi K2.511.5%estimated ± 12.0 pp, low confidence
89Kimi K2.5 (Reasoning)11.5%estimated ± 12.0 pp, low confidence
90GPT-5.111.0%estimated ± 12.0 pp, low confidence
91Qwen3.5-122B-A10B10.5%estimated ± 12.0 pp, low confidence
92MiniMax M310.0%measured
93Kimi K2.610.0%estimated ± 5.5 pp, medium confidence
94Inkling-Small9.8%estimated ± 6.7 pp, low confidence
95Gemini 3.1 Pro9.7%estimated ± 12.0 pp, low confidence
96Qwen3.5 397B9.5%estimated ± 12.5 pp, low confidence
97Qwen3.7 Plus9.2%estimated ± 5.5 pp, low confidence
98Qwen3.5-27B9.0%estimated ± 12.5 pp, low confidence
99Qwen3.5-35B-A3B9.0%estimated ± 12.5 pp, low confidence
100Mistral Medium 3.5 128B8.6%estimated ± 12.0 pp, low confidence
101MiniCPM5-2B6.8%estimated ± 12.0 pp, low confidence
102Nemotron 3 Ultra6.3%measured
103Laguna S 2.15.9%estimated ± 6.7 pp, low confidence
104LongCat-Flash-Lite-Sparse5.0%estimated ± 12.5 pp, low confidence
105Ornith-1.5-35B-A3B4.9%estimated ± 2.5 pp, low confidence
106Ornith-1.5-9B4.5%estimated ± 2.5 pp, low confidence
107Nemotron 3.5 Lightning 30B A3B NVFP44.4%estimated ± 12.0 pp, low confidence
108MiMo-V2-Flash4.2%estimated ± 12.0 pp, low confidence
109Gemma 4 31B3.8%estimated ± 12.0 pp, low confidence
110GPT-OSS 120B3.5%estimated ± 12.0 pp, low confidence
111Granite 4.2 30B2.4%estimated ± 12.0 pp, low confidence
112Gemma 4 26B A4B2.1%estimated ± 12.0 pp, low confidence
113Ling 3.0 Tiny2.0%estimated ± 12.0 pp, low confidence
114Command A+0.6%estimated ± 12.0 pp, low confidence
115Celeris-10.0%estimated ± 12.0 pp, low confidence
116DeepSeek V30.0%estimated ± 12.0 pp, low confidence
117DeepSeek V3 03240.0%estimated ± 12.0 pp, low confidence
118Gemini 2.5 Pro0.0%estimated ± 12.0 pp, low confidence
119Gemma 3 27B0.0%estimated ± 12.0 pp, low confidence
120Gemma 4 12B0.0%estimated ± 12.0 pp, low confidence
121Gemma 4 E2B0.0%estimated ± 12.0 pp, low confidence
122Gemma 4 E4B0.0%estimated ± 12.0 pp, low confidence
123GPT-4.1 mini0.0%estimated ± 12.0 pp, low confidence
124GPT-4.1 nano0.0%estimated ± 12.0 pp, low confidence
125GPT-4o0.0%estimated ± 12.0 pp, low confidence
126GPT-4o mini0.0%estimated ± 12.0 pp, low confidence
127GPT-OSS 20B0.0%estimated ± 12.0 pp, low confidence
128Granite 4.2 3B0.0%estimated ± 12.0 pp, low confidence
129Granite 4.2 8B0.0%estimated ± 12.0 pp, low confidence
130K-Exaone0.0%estimated ± 12.0 pp, low confidence
131LFM2.5-2.6B0.0%estimated ± 12.0 pp, low confidence
132Ling 2.6 Flash0.0%estimated ± 12.0 pp, low confidence
133Llama 4 Maverick0.0%estimated ± 12.0 pp, low confidence
134Llama 4 Scout0.0%estimated ± 12.0 pp, low confidence
135Mercury 2.50.0%estimated ± 12.0 pp, low confidence
136Mistral Large 30.0%estimated ± 12.0 pp, low confidence
137Mistral Small 40.0%estimated ± 12.0 pp, low confidence
138Mistral Small 4 (Reasoning)0.0%estimated ± 12.0 pp, low confidence
139Nemotron 3 Nano 30B0.0%estimated ± 12.0 pp, low confidence
140Nemotron 3 Nano Omni 30B A3B0.0%estimated ± 12.0 pp, low confidence
141Nemotron 3 Super 100B0.0%estimated ± 12.0 pp, low confidence
142North Mini Code0.0%estimated ± 12.0 pp, low confidence
143Solar Pro 30.0%estimated ± 12.0 pp, low confidence
144Trinity-Large-Preview0.0%estimated ± 12.0 pp, low confidence
145Trinity-Large-Thinking0.0%estimated ± 12.0 pp, low confidence
146Ultravox v0.6 Llama 3.3 70B0.0%estimated ± 12.0 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General