benchgap
Agentic · tools

τ³-bench results leaderboard

As of 2026-10-07, the highest measured score on τ³-bench results is 96.0% by Mercury 2.5. 74 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1DeepSeek V3.2100.0%estimated ± 8.9 pp, low confidence
2GLM-5V-Turbo100.0%estimated ± 8.9 pp, low confidence
3GPT-4.1100.0%estimated ± 8.9 pp, low confidence
4GPT-OSS 120B100.0%estimated ± 8.9 pp, low confidence
5Kimi K2.5 (Reasoning)100.0%estimated ± 8.9 pp, low confidence
6Qwen3.5-35B-A3B100.0%estimated ± 8.9 pp, low confidence
7Trinity-Large-Thinking100.0%estimated ± 8.9 pp, low confidence
8Gemma 4 31B100.0%estimated ± 8.9 pp, low confidence
9MiMo-V2-Pro99.9%estimated ± 8.9 pp, low confidence
10Hy3 Preview99.9%estimated ± 8.9 pp, low confidence
11Grok 4.2097.5%estimated ± 8.9 pp, low confidence
12Gemini 3.1 Flash-Lite97.0%estimated ± 8.9 pp, low confidence
13Mercury 2.596.0%measured
14Mistral Medium 3.5 128B91.4%measured
15Qwen3.5-27B87.1%estimated ± 8.9 pp, medium confidence
16Claude 4 Sonnet83.2%estimated ± 8.9 pp, medium confidence
17GLM-4.778.9%estimated ± 8.9 pp, medium confidence
18MiMo-V2.5-Pro72.9%measured
19GPT-5.6 Sol72.9%estimated ± 2.1 pp, low confidence
20GPT-5.6 Luna71.9%estimated ± 2.1 pp, low confidence
21GPT-5.6 Terra71.8%estimated ± 2.1 pp, low confidence
22Step 3.7 Flash71.0%estimated ± 2.1 pp, low confidence
23Nemotron 3 Ultra70.9%measured
24Qwen3.6 Plus70.7%measured
25Qwen3.7 Plus70.7%estimated ± 1.4 pp, low confidence
26GLM-5.170.6%measured
27MiniMax M2.770.3%estimated ± 2.1 pp, low confidence
28Claude Opus 4.570.2%measured
29GPT-5.170.0%estimated ± 8.9 pp, medium confidence
30Muse Spark 1.169.6%estimated ± 1.8 pp, low confidence
31Claude Opus 569.6%estimated ± 1.8 pp, low confidence
32Step 5 Preview69.6%estimated ± 1.8 pp, low confidence
33Kimi K369.6%estimated ± 1.8 pp, low confidence
34Gemini 3.5 Flash69.6%estimated ± 1.8 pp, low confidence
35Hy4 preview69.6%estimated ± 1.8 pp, low confidence
36Claude Opus 4.869.6%estimated ± 1.8 pp, low confidence
37Ornith-1.5-397B69.6%estimated ± 1.8 pp, low confidence
38Inkling-Small69.6%estimated ± 1.8 pp, low confidence
39Beam69.6%estimated ± 1.8 pp, low confidence
40Claude Opus 4.7 (Adaptive)69.6%estimated ± 1.8 pp, low confidence
41GLM-5.269.6%estimated ± 1.8 pp, low confidence
42Qwen3.7 Max69.6%estimated ± 1.8 pp, low confidence
43Kimi K2.7 Code69.6%estimated ± 1.8 pp, low confidence
44Muse Glimmer 30B69.6%estimated ± 1.8 pp, low confidence
45GPT-5.569.6%estimated ± 1.8 pp, low confidence
46MiniMax M369.6%estimated ± 1.8 pp, low confidence
47Inkling69.6%estimated ± 1.8 pp, low confidence
48DeepSeek V4 Pro 081369.6%estimated ± 1.8 pp, low confidence
49GPT-5.469.6%estimated ± 1.8 pp, medium confidence
50Ornith-1.5-35B-A3B69.6%estimated ± 1.8 pp, medium confidence
51DeepSeek V4 Flash 073169.6%estimated ± 1.8 pp, medium confidence
52Ling 3.0 Flash69.6%estimated ± 1.8 pp, medium confidence
53Solar Pro 469.6%estimated ± 1.8 pp, medium confidence
54Solar Open 269.5%estimated ± 1.8 pp, medium confidence
55GPT-5.4 mini69.5%estimated ± 1.8 pp, medium confidence
56GPT-5.4 nano69.5%estimated ± 1.8 pp, medium confidence
57Kimi K2.669.5%estimated ± 1.8 pp, medium confidence
58Ornith-1.5-9B69.5%estimated ± 1.8 pp, medium confidence
59LLaDA2.2-flash69.3%estimated ± 1.8 pp, medium confidence
60LongCat-Flash-Lite-Sparse69.2%estimated ± 1.8 pp, medium confidence
61Gemini 2.5 Pro69.1%estimated ± 8.9 pp, medium confidence
62Grok 469.0%estimated ± 8.9 pp, medium confidence
63Qwen3 Max68.9%estimated ± 8.9 pp, medium confidence
64Grok 4.368.9%estimated ± 8.9 pp, medium confidence
65Claude Opus 4.668.9%estimated ± 8.9 pp, medium confidence
66Claude Opus 4.768.9%estimated ± 8.9 pp, low confidence
67Claude Sonnet 4.568.9%estimated ± 8.9 pp, medium confidence
68Claude Sonnet 4.668.9%estimated ± 8.9 pp, medium confidence
69Gemini 3.1 Pro68.9%estimated ± 8.9 pp, medium confidence
70Gemini 3 Flash68.9%estimated ± 8.9 pp, medium confidence
71Gemini 3 Pro68.9%estimated ± 8.9 pp, medium confidence
72GPT-5.1-Codex68.9%estimated ± 8.9 pp, medium confidence
73GPT-5.268.9%estimated ± 8.9 pp, medium confidence
74GPT-5.2-Codex68.9%estimated ± 8.9 pp, medium confidence
75GPT-5.3 Codex68.9%estimated ± 8.9 pp, medium confidence
76Grok 4.1 Fast68.9%estimated ± 8.9 pp, medium confidence
77Grok Build 0.168.9%estimated ± 8.9 pp, medium confidence
78MiMo-V2.568.9%estimated ± 8.9 pp, medium confidence
79Qwen3.6-27B68.9%estimated ± 8.9 pp, medium confidence
80Qwen3.5 397B68.4%measured
81dots3-note Preview67.9%estimated ± 11.4 pp, low confidence
82Qwen3.6-35B-A3B67.2%measured
83Pokee-Isaac 28B66.2%measured
84Kimi K2.565.7%measured
85GLM-565.6%measured
86Granite 4.2 30B62.0%measured
87Granite 4.2 8B58.1%measured
88Granite 4.2 3B45.8%measured
89Qwen3.8 Max41.2%estimated ± 11.4 pp, low confidence
90Atria Dawn Preview41.2%measured
91Nemotron 3.5 Lightning 30B A3B NVFP49.5%measured
92LFM2.5-2.6B5.7%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General