benchgap
Agentic · tools

AA Tau3 Banking leaderboard

As of 2026-10-07, the highest measured score on AA Tau3 Banking is 50.7% by Grok 4.6. 135 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Grok 4.650.7%measured
2Muse Spark 1.350.5%measured
3GLM-5.350.3%measured
4Claude Opus 5.549.3%estimated ± 3.0 pp, low confidence
5Grok 4.748.2%estimated ± 3.0 pp, low confidence
6GLM-5.248.0%estimated ± 3.0 pp, low confidence
7Qwen3.8-27B48.0%measured
8Gemini 3.7 Flash47.7%estimated ± 4.4 pp, low confidence
9DeepSeek V4 Pro 081347.5%estimated ± 5.8 pp, medium confidence
10Gemini 4 Argon47.5%estimated ± 5.8 pp, medium confidence
11Ling 3.1 Flash47.5%estimated ± 5.8 pp, medium confidence
12MiMo-V2.6-Flash47.5%estimated ± 5.8 pp, medium confidence
13MiMo-V2.6-Pro47.5%estimated ± 5.8 pp, medium confidence
14Qwen3.8-Flash-Next47.5%estimated ± 5.8 pp, medium confidence
15Qwen3.8 Max Preview47.5%estimated ± 5.8 pp, medium confidence
16Muse Spark 1.247.5%estimated ± 5.8 pp, medium confidence
17Hy4 preview47.4%estimated ± 10.7 pp, low confidence
18Claude Sonnet 547.4%estimated ± 5.8 pp, medium confidence
19GPT-5.6 Luna47.4%estimated ± 5.8 pp, medium confidence
20Claude Opus 4.847.4%estimated ± 5.8 pp, medium confidence
21DeepSeek V4 Flash 073147.2%estimated ± 5.8 pp, medium confidence
22GPT-6 Luna47.2%estimated ± 5.8 pp, medium confidence
23Claude Fable 5.147.2%measured
24GLM-5.3-Flash47.2%measured
25GPT-5.547.1%estimated ± 3.0 pp, low confidence
26Claude Sonnet 5.547.0%estimated ± 4.4 pp, medium confidence
27Mistral Large 446.9%estimated ± 5.8 pp, medium confidence
28Claude Opus 4.7 (Adaptive)46.9%estimated ± 3.0 pp, low confidence
29DeepSeek V4.1 Flash46.8%estimated ± 3.0 pp, low confidence
30GPT-6 Sol46.1%estimated ± 3.0 pp, low confidence
31Kimi K346.0%measured
32Claude Opus 545.9%estimated ± 4.4 pp, medium confidence
33GPT-5.6 Terra45.6%estimated ± 3.0 pp, low confidence
34Gemini 3.8 Flash44.9%measured
35GPT-6.1 Sol44.6%estimated ± 4.4 pp, medium confidence
36Step 5 Preview44.3%estimated ± 3.0 pp, low confidence
37Claude Fable 544.2%estimated ± 4.4 pp, medium confidence
38GPT-5.6 Sol44.1%estimated ± 3.0 pp, low confidence
39Grok 4.542.5%estimated ± 5.8 pp, medium confidence
40Atria Dawn Preview41.4%estimated ± 10.2 pp, low confidence
41GPT-6 Astra41.4%measured
42GPT-5.5 Pro38.4%estimated ± 10.2 pp, low confidence
43GPT-5.4 Pro37.4%estimated ± 10.2 pp, low confidence
44Claude Mythos 535.9%estimated ± 10.2 pp, low confidence
45Qwen3.8 Max35.3%estimated ± 14.6 pp, low confidence
46Claude Haiku 5.534.8%estimated ± 7.1 pp, medium confidence
47Ornith-1.5-397B34.3%estimated ± 10.2 pp, low confidence
48Claude Opus 4.631.3%estimated ± 10.2 pp, low confidence
49Claude Opus 4.731.0%estimated ± 14.6 pp, low confidence
50dots3-note Preview30.9%estimated ± 10.2 pp, low confidence
51Inkling29.1%measured
52Ling 3.0 Flash28.0%measured
53Gemini 3.5 Flash26.0%estimated ± 5.8 pp, medium confidence
54Beam25.7%estimated ± 10.2 pp, low confidence
55Agents-A124.3%estimated ± 10.2 pp, low confidence
56Claude Sonnet 4.624.1%estimated ± 14.6 pp, low confidence
57Muse Glimmer 30B23.5%measured
58Gemini 3.6 Flash22.1%estimated ± 5.8 pp, medium confidence
59GPT-5.422.0%estimated ± 5.8 pp, medium confidence
60Hy3 Preview22.0%estimated ± 5.8 pp, medium confidence
61Muse Spark 1.122.0%estimated ± 5.8 pp, medium confidence
62Apodex 1.122.0%estimated ± 5.8 pp, medium confidence
63Apodex 1.1 Mini22.0%estimated ± 5.8 pp, medium confidence
64Quasar 438B22.0%estimated ± 5.8 pp, medium confidence
65A.X K222.0%estimated ± 5.8 pp, medium confidence
66Celeris-122.0%estimated ± 5.8 pp, low confidence
67Command A+22.0%estimated ± 5.8 pp, low confidence
68DeepSeek V322.0%estimated ± 5.8 pp, low confidence
69DeepSeek V3 032422.0%estimated ± 5.8 pp, low confidence
70Gemini 2.5 Pro22.0%estimated ± 5.8 pp, low confidence
71Gemini 3.1 Pro22.0%estimated ± 5.8 pp, medium confidence
72Gemini 3.5 Flash-Lite22.0%estimated ± 5.8 pp, medium confidence
73Gemma 3 27B22.0%estimated ± 5.8 pp, low confidence
74Gemma 4 12B22.0%estimated ± 5.8 pp, low confidence
75Gemma 4 26B A4B22.0%estimated ± 5.8 pp, low confidence
76Gemma 4 31B22.0%estimated ± 5.8 pp, low confidence
77Gemma 4 E2B22.0%estimated ± 5.8 pp, low confidence
78Gemma 4 E4B22.0%estimated ± 5.8 pp, low confidence
79GLM-4.722.0%estimated ± 5.8 pp, medium confidence
80GLM-5.122.0%estimated ± 5.8 pp, medium confidence
81GPT-4.1 mini22.0%estimated ± 5.8 pp, low confidence
82GPT-4.1 nano22.0%estimated ± 5.8 pp, low confidence
83GPT-4o22.0%estimated ± 5.8 pp, low confidence
84GPT-4o mini22.0%estimated ± 5.8 pp, low confidence
85GPT-5.122.0%estimated ± 5.8 pp, medium confidence
86GPT-5.4 mini22.0%estimated ± 5.8 pp, medium confidence
87GPT-5.4 nano22.0%estimated ± 5.8 pp, medium confidence
88GPT-5 (high)22.0%estimated ± 5.8 pp, medium confidence
89GPT-OSS 120B22.0%estimated ± 5.8 pp, low confidence
90GPT-OSS 20B22.0%estimated ± 5.8 pp, low confidence
91Granite 4.2 30B22.0%estimated ± 5.8 pp, low confidence
92Granite 4.2 3B22.0%estimated ± 5.8 pp, low confidence
93Granite 4.2 8B22.0%estimated ± 5.8 pp, low confidence
94Grok 4.322.0%estimated ± 5.8 pp, medium confidence
95Hy322.0%estimated ± 5.8 pp, medium confidence
96Inkling-Small22.0%estimated ± 5.8 pp, medium confidence
97K-Exaone22.0%estimated ± 5.8 pp, low confidence
98Kimi K2.622.0%estimated ± 5.8 pp, medium confidence
99Kimi K2.522.0%estimated ± 5.8 pp, medium confidence
100Kimi K2.5 (Reasoning)22.0%estimated ± 5.8 pp, medium confidence
101Kimi K2.7 Code22.0%estimated ± 5.8 pp, medium confidence
102LFM2.5-2.6B22.0%estimated ± 5.8 pp, low confidence
103Ling 2.6 Flash22.0%estimated ± 5.8 pp, low confidence
104Ling 3.0 Flash FP822.0%estimated ± 5.8 pp, medium confidence
105Ling 3.0 Flash VL22.0%estimated ± 5.8 pp, medium confidence
106Ling 3.0 Tiny22.0%estimated ± 5.8 pp, low confidence
107Llama 4 Maverick22.0%estimated ± 5.8 pp, low confidence
108Llama 4 Scout22.0%estimated ± 5.8 pp, low confidence
109Mercury 2.522.0%estimated ± 5.8 pp, low confidence
110MiMo-V2.5-Pro22.0%estimated ± 5.8 pp, medium confidence
111MiMo-V2-Flash22.0%estimated ± 5.8 pp, low confidence
112MiniCPM5-2B22.0%estimated ± 5.8 pp, low confidence
113MiniMax M2.722.0%estimated ± 5.8 pp, medium confidence
114Mistral Large 322.0%estimated ± 5.8 pp, low confidence
115Mistral Medium 3.5 128B22.0%estimated ± 5.8 pp, low confidence
116Mistral Small 422.0%estimated ± 5.8 pp, low confidence
117Mistral Small 4 (Reasoning)22.0%estimated ± 5.8 pp, low confidence
118Muse Spark22.0%estimated ± 5.8 pp, medium confidence
119Nemotron 3.5 Lightning 30B A3B NVFP422.0%estimated ± 5.8 pp, low confidence
120Nemotron 3 Nano 30B22.0%estimated ± 5.8 pp, low confidence
121Nemotron 3 Nano Omni 30B A3B22.0%estimated ± 5.8 pp, low confidence
122Nemotron 3 Super 100B22.0%estimated ± 5.8 pp, low confidence
123North Mini Code22.0%estimated ± 5.8 pp, low confidence
124Qwen3.5-122B-A10B22.0%estimated ± 5.8 pp, medium confidence
125Qwen3.6-27B22.0%estimated ± 5.8 pp, medium confidence
126Qwen3.6-35B-A3B22.0%estimated ± 5.8 pp, medium confidence
127Qwen3.6 Plus22.0%estimated ± 5.8 pp, medium confidence
128Qwen3.7 Max22.0%estimated ± 5.8 pp, medium confidence
129Qwen3.7 Plus22.0%estimated ± 5.8 pp, low confidence
130Solar Pro 322.0%estimated ± 5.8 pp, low confidence
131Solar Pro 422.0%estimated ± 5.8 pp, medium confidence
132Step 3.7 Flash22.0%estimated ± 5.8 pp, medium confidence
133Trinity-Large-Preview22.0%estimated ± 5.8 pp, low confidence
134Trinity-Large-Thinking22.0%estimated ± 5.8 pp, low confidence
135Ultravox v0.6 Llama 3.3 70B22.0%estimated ± 5.8 pp, low confidence
136Ornith-1.5-35B-A3B19.9%estimated ± 10.2 pp, low confidence
137Agents-A1-4B19.6%estimated ± 10.2 pp, low confidence
138GPT-5.219.2%estimated ± 10.2 pp, low confidence
139Qwen3.5 397B17.9%estimated ± 10.2 pp, low confidence
140Qwen3.5-27B17.6%estimated ± 10.2 pp, low confidence
141Qwen3.5-35B-A3B17.6%estimated ± 10.2 pp, low confidence
142Ornith-1.5-9B16.5%estimated ± 10.2 pp, low confidence
143LongCat-Flash-Lite-Sparse15.5%estimated ± 10.2 pp, low confidence
144MiniMax M315.3%measured
145Nemotron 3 Ultra14.2%measured
146Solar Open 213.0%estimated ± 11.8 pp, low confidence
147LLaDA2.2-flash8.2%estimated ± 11.8 pp, low confidence
148Claude Opus 4.57.0%estimated ± 11.8 pp, low confidence
149GLM-54.4%estimated ± 11.8 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General