Agentic · tools
τ³-Bench Banking leaderboard
As of 2026-10-07, the highest measured score on τ³-Bench Banking is 51.3% by Qwen3.8 Max. 28 more models have estimated scores, calibrated from the benchmarks they were measured on.
Measured scores: artificial-analysis.
| # | Model | Score | Source |
|---|---|---|---|
| 1 | Qwen3.8 Max | 51.3% | measured |
| 2 | Grok 4.6 (high) | 50.7% | measured |
| 3 | Muse Spark 1.3 (max) | 50.5% | measured |
| 4 | Gemini 3.7 Flash (medium) | 50.4% | estimated ± 11.4 pp, low confidence |
| 5 | GLM-5.3 (max) | 50.3% | measured |
| 6 | Gemini 3.5 Flash (high) | 50.3% | estimated ± 11.4 pp, low confidence |
| 7 | DeepSeek V4 Pro 0813 (max) | 50.1% | estimated ± 11.4 pp, low confidence |
| 8 | Qwen3.8 2.4T A95B | 49.1% | measured |
| 9 | Claude Opus 5.5 (max with fallback) | 49.0% | estimated ± 2.6 pp, medium confidence |
| 10 | Gemini 3.7 Flash (high) | 49.0% | estimated ± 4.5 pp, low confidence |
| 11 | Claude Sonnet 5.5 (max with fallback) | 48.3% | estimated ± 4.5 pp, medium confidence |
| 12 | Qwen3.8 27B (xhigh) | 48.0% | measured |
| 13 | Grok 4.7 (xhigh) | 48.0% | estimated ± 2.6 pp, medium confidence |
| 14 | GLM-5.2 (max) | 47.8% | estimated ± 2.6 pp, medium confidence |
| 15 | Qwen3.8 Max (0902) | 47.8% | measured |
| 16 | GPT-6 Astra (high) | 47.4% | estimated ± 5.0 pp, medium confidence |
| 17 | GPT-6 Astra (xhigh) | 47.4% | estimated ± 5.0 pp, medium confidence |
| 18 | Gemini 4 Argon (high) | 47.4% | estimated ± 5.0 pp, low confidence |
| 19 | MiMo-V2.6-Pro | 47.4% | estimated ± 5.0 pp, medium confidence |
| 20 | Mistral Large 4 Preview | 47.4% | estimated ± 5.0 pp, medium confidence |
| 21 | GPT-6 Luna (max) | 47.4% | estimated ± 5.0 pp, medium confidence |
| 22 | GPT-6.1 Sol (xhigh) | 47.4% | estimated ± 5.0 pp, medium confidence |
| 23 | Qwen3.8 27B (medium) | 47.4% | measured |
| 24 | Grok 4.6 (xhigh) | 47.4% | estimated ± 5.0 pp, medium confidence |
| 25 | Claude Fable 5.1 (max with fallback) | 47.2% | measured |
| 26 | GLM-5.3-Flash | 47.2% | measured |
| 27 | Muse Spark 1.3 (xhigh) | 47.2% | measured |
| 28 | Claude Opus 5 (max) | 47.1% | estimated ± 4.5 pp, medium confidence |
| 29 | GPT-5.5 (xhigh) | 47.0% | estimated ± 2.6 pp, medium confidence |
| 30 | Claude Opus 4.7 (max) | 46.8% | estimated ± 2.6 pp, medium confidence |
| 31 | DeepSeek V4.1 Flash (max) | 46.7% | estimated ± 2.6 pp, medium confidence |
| 32 | GPT-6 Sol (max) | 46.0% | estimated ± 2.6 pp, medium confidence |
| 33 | Kimi K3 (max) | 46.0% | measured |
| 34 | GPT-6.1 Sol (max) | 45.8% | estimated ± 4.5 pp, medium confidence |
| 35 | GPT-5.6 Terra (max) | 45.6% | estimated ± 2.6 pp, medium confidence |
| 36 | Claude Fable 5 (with fallback) | 45.3% | estimated ± 4.5 pp, medium confidence |
| 37 | Gemini 3.8 Flash (high) | 44.9% | measured |
| 38 | Step 5 Preview | 44.4% | estimated ± 2.6 pp, low confidence |
| 39 | GPT-5.6 Sol (max) | 44.2% | estimated ± 2.6 pp, low confidence |
| 40 | GPT-6 Astra (max) | 41.4% | measured |
| 41 | Claude Fable 5.1 (xhigh with fallback) | 40.5% | estimated ± 10.4 pp, low confidence |
| 42 | Claude Fable 5.1 (high with fallback) | 39.9% | estimated ± 10.4 pp, low confidence |
| 43 | K2 Horizon 375B A23B | 34.2% | measured |
| 44 | Inkling (xhigh) | 29.1% | measured |
| 45 | Muse Glimmer (high) | 23.5% | measured |
| 46 | MiniMax-M3 | 15.3% | measured |
| 47 | Nemotron 3 Ultra | 14.2% | measured |
All leaderboards
Agentic · terminal
Agentic · tools
- AA-AnalystAgent
- APEX-Agents
- AutomationBench
- EnterpriseOps-Gym
- Harvey LAB
- ITBench SRE
- τ³-Bench Banking
- τ²-Bench Telecom