benchgap
Agentic · tools

AA-AnalystAgent leaderboard

As of 2026-10-07, the highest measured score on AA-AnalystAgent is 60.0% by Gemini 3.7 Flash (high). 28 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: artificial-analysis.

#ModelScoreSource
1Gemini 3.7 Flash (high)60.0%measured
2Gemini 4 Argon (high)58.3%estimated ± 7.8 pp, low confidence
3Claude Fable 5.1 (max with fallback)57.5%measured
4Claude Sonnet 5.5 (max with fallback)57.5%measured
5Claude Opus 5.5 (max with fallback)56.3%measured
6Qwen3.8 Max54.8%estimated ± 8.1 pp, low confidence
7DeepSeek V4.1 Flash (max)54.8%estimated ± 7.8 pp, medium confidence
8Claude Opus 5 (max)53.8%measured
9GPT-6 Astra (xhigh)53.8%estimated ± 7.8 pp, medium confidence
10Grok 4.6 (xhigh)53.6%estimated ± 7.8 pp, medium confidence
11GPT-6 Astra (high)53.4%estimated ± 7.8 pp, medium confidence
12GPT-6.1 Sol (xhigh)53.4%estimated ± 7.8 pp, medium confidence
13Grok 4.7 (xhigh)52.7%estimated ± 7.8 pp, medium confidence
14Qwen3.8 2.4T A95B52.0%estimated ± 8.1 pp, low confidence
15GPT-6 Astra (max)51.2%measured
16GPT-6.1 Sol (max)50.0%measured
17GPT-5.5 (xhigh)50.0%measured
18Qwen3.8 27B (medium)49.8%estimated ± 8.1 pp, low confidence
19Muse Spark 1.3 (xhigh)49.5%estimated ± 8.1 pp, low confidence
20GPT-6 Sol (max)49.4%estimated ± 7.8 pp, medium confidence
21Claude Fable 5 (with fallback)48.8%measured
22Gemini 3.8 Flash (high)47.8%estimated ± 7.8 pp, medium confidence
23Mistral Large 4 Preview47.8%estimated ± 7.8 pp, medium confidence
24GPT-5.6 Sol (max)47.5%measured
25Gemini 3.7 Flash (medium)47.5%estimated ± 3.6 pp, medium confidence
26Claude Fable 5.1 (xhigh with fallback)47.3%estimated ± 13.9 pp, low confidence
27Gemini 3.5 Flash (high)46.8%estimated ± 3.6 pp, medium confidence
28MiMo-V2.6-Pro46.4%estimated ± 7.8 pp, medium confidence
29Claude Fable 5.1 (high with fallback)46.1%estimated ± 13.9 pp, low confidence
30DeepSeek V4 Pro 0813 (max)45.7%estimated ± 3.6 pp, medium confidence
31Muse Spark 1.3 (max)45.6%estimated ± 7.8 pp, medium confidence
32Qwen3.8 Max (0902)45.0%measured
33Grok 4.6 (high)43.0%estimated ± 3.6 pp, medium confidence
34GPT-6 Luna (max)39.6%estimated ± 7.8 pp, medium confidence
35Kimi K3 (max)38.8%measured
36Step 5 Preview35.0%measured
37Qwen3.8 27B (xhigh)34.5%estimated ± 3.6 pp, medium confidence
38Inkling (xhigh)23.8%measured
39GLM-5.3 (max)19.4%estimated ± 3.6 pp, medium confidence
40K2 Horizon 375B A23B19.1%estimated ± 7.8 pp, medium confidence
41Muse Glimmer (high)16.2%estimated ± 3.6 pp, medium confidence
42GLM-5.3-Flash13.5%estimated ± 3.6 pp, medium confidence
43MiniMax-M310.0%measured
44Nemotron 3 Ultra6.3%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Vision & documents