benchgap
Agentic · tools

AA ITBench leaderboard

As of 2026-10-07, the highest measured score on AA ITBench is 56.2% by GPT-5.6 Sol. 159 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Exaone 4.0 32B100.0%estimated ± 3.6 pp, low confidence
2Phi-4100.0%estimated ± 3.6 pp, low confidence
3LFM2.5-VL-1.6B-Extract100.0%estimated ± 3.6 pp, low confidence
4Gemma 3 27B100.0%estimated ± 3.6 pp, low confidence
5Nemotron Ultra 253B100.0%estimated ± 3.6 pp, low confidence
6Granite-4.0-350M100.0%estimated ± 3.6 pp, low confidence
7Nova Pro100.0%estimated ± 3.6 pp, low confidence
8Granite-4.0-H-350M100.0%estimated ± 3.6 pp, low confidence
9Gemini 2.5 Flash100.0%estimated ± 3.6 pp, low confidence
10Llama 4 Scout100.0%estimated ± 3.6 pp, low confidence
11Qwen3-Omni-30B-A3B-Instruct100.0%estimated ± 3.6 pp, low confidence
12GPT-4.1 nano100.0%estimated ± 3.6 pp, low confidence
13Llama 4 Maverick100.0%estimated ± 3.6 pp, low confidence
14Llama 3.1 405B100.0%estimated ± 3.6 pp, low confidence
15Granite-4.0-H-1B100.0%estimated ± 3.6 pp, low confidence
16Exaone 4.0 1.2B100.0%estimated ± 3.6 pp, low confidence
17Gemma 4 E2B100.0%estimated ± 3.6 pp, low confidence
18Gemma 4 E4B100.0%estimated ± 3.6 pp, low confidence
19Claude 3 Haiku100.0%estimated ± 3.6 pp, low confidence
20Qwen3-Omni-30B-A3B-Thinking100.0%estimated ± 3.6 pp, low confidence
21DeepSeek V399.9%estimated ± 3.6 pp, low confidence
22Mistral Medium 399.9%estimated ± 3.6 pp, low confidence
23Mistral Large 399.9%estimated ± 3.6 pp, low confidence
24GPT-4o99.9%estimated ± 3.6 pp, low confidence
25Ultravox v0.6 Llama 3.3 70B99.8%estimated ± 3.6 pp, low confidence
26o3-mini99.8%estimated ± 3.6 pp, low confidence
27Mistral Large 299.6%estimated ± 3.6 pp, low confidence
28Solar Pro 299.5%estimated ± 3.6 pp, low confidence
29Sarvam 30B99.3%estimated ± 3.6 pp, low confidence
30DeepSeek V3.199.2%estimated ± 3.6 pp, low confidence
31Gemma 4 12B99.0%estimated ± 3.6 pp, low confidence
32DeepSeek-R199.0%estimated ± 3.6 pp, low confidence
33DeepSeek V3.1 (Reasoning)98.8%estimated ± 3.6 pp, low confidence
34North Mini Code98.8%estimated ± 3.6 pp, low confidence
35Nemotron 3 Nano 30B98.0%estimated ± 3.6 pp, low confidence
36Mistral Small 497.9%estimated ± 3.6 pp, low confidence
37Mistral Small 4 (Reasoning)97.9%estimated ± 3.6 pp, low confidence
38Gemini 3 Flash97.3%estimated ± 3.6 pp, low confidence
39Gemma 4 26B A4B97.1%estimated ± 3.6 pp, low confidence
40Nemotron 3 Nano Omni 30B A3B96.4%estimated ± 3.6 pp, low confidence
41GLM-4.5-Air95.9%estimated ± 3.6 pp, low confidence
42Sarvam 105B95.7%estimated ± 3.6 pp, low confidence
43DeepSeek V3 032495.6%estimated ± 3.6 pp, low confidence
44GPT-4.195.6%estimated ± 3.6 pp, low confidence
45Claude 4 Sonnet92.2%estimated ± 3.6 pp, low confidence
46GPT-4.1 mini91.7%estimated ± 3.6 pp, low confidence
47Gemini 2.5 Pro90.7%estimated ± 3.6 pp, low confidence
48LLaDA2.2-mini87.3%estimated ± 3.6 pp, low confidence
49Gemma 4 31B84.6%estimated ± 3.6 pp, low confidence
50GPT-OSS 20B84.3%estimated ± 3.6 pp, low confidence
51Kimi K283.2%estimated ± 3.6 pp, low confidence
52o181.3%estimated ± 3.6 pp, low confidence
53Grok 4.1 Fast79.9%estimated ± 3.6 pp, low confidence
54GPT-OSS 120B77.1%estimated ± 3.6 pp, low confidence
55Grok 4 Fast (Reasoning)77.1%estimated ± 3.6 pp, low confidence
56Nemotron 3 Super 100B74.3%estimated ± 3.6 pp, low confidence
57Ornith-1.5-9B73.2%estimated ± 3.4 pp, low confidence
58Ornith-1.5-35B-A3B70.7%estimated ± 3.4 pp, low confidence
59Claude 4.1 Opus Thinking69.3%estimated ± 3.6 pp, low confidence
60Nemotron 3 Ultra67.2%estimated ± 3.4 pp, low confidence
61Claude Opus 4.765.8%estimated ± 3.6 pp, low confidence
62K-Exaone65.4%estimated ± 3.6 pp, low confidence
63Qwen3 Max65.4%estimated ± 3.6 pp, low confidence
64Grok 464.6%estimated ± 3.6 pp, low confidence
65Grok Code Fast 163.6%estimated ± 3.6 pp, low confidence
66GLM-4.662.1%estimated ± 3.6 pp, low confidence
67Agents-A1-4B60.5%estimated ± 3.6 pp, low confidence
68DeepSeek V4 Flash 073160.4%estimated ± 3.4 pp, low confidence
69DeepSeek V3.259.6%estimated ± 3.6 pp, low confidence
70Claude Sonnet 4.658.9%estimated ± 3.6 pp, low confidence
71Step 3.7 Flash58.6%estimated ± 3.4 pp, low confidence
72Agents-A158.2%estimated ± 3.4 pp, low confidence
73Ternary Bonsai 2 27B58.1%estimated ± 3.6 pp, low confidence
74o357.6%estimated ± 3.6 pp, low confidence
75GPT-5.156.3%estimated ± 3.6 pp, low confidence
76Muse Spark 1.156.3%estimated ± 3.6 pp, low confidence
77GPT-5.6 Sol56.2%measured
78Step 5 Preview55.6%measured
79GPT-5.1-Codex55.1%estimated ± 3.6 pp, low confidence
80GPT-5.1-Codex-Max55.1%estimated ± 3.6 pp, low confidence
81MiMo-V2-Flash54.2%estimated ± 3.6 pp, low confidence
82Qwen3.5 397B (Reasoning)54.2%estimated ± 3.6 pp, low confidence
83dots3-note Preview53.8%estimated ± 3.4 pp, low confidence
84Claude Opus 4.653.3%estimated ± 3.6 pp, low confidence
85GPT-5.253.3%estimated ± 3.6 pp, low confidence
86GPT-5 (high)53.3%estimated ± 3.6 pp, low confidence
87MiniMax M2.753.3%estimated ± 3.6 pp, low confidence
88Command A+53.1%estimated ± 3.6 pp, low confidence
89Qwen3.7 Max53.1%estimated ± 3.4 pp, low confidence
90Gemini 3.8 Flash52.5%measured
91GPT-6.1 Sol52.4%estimated ± 3.9 pp, medium confidence
92GPT-5.3 Codex52.2%estimated ± 3.6 pp, medium confidence
93Ling 2.6 Flash52.2%estimated ± 3.6 pp, medium confidence
94Solar Pro 352.0%estimated ± 3.6 pp, medium confidence
95GPT-5 (medium)51.8%estimated ± 3.6 pp, medium confidence
96Hy4 preview51.4%estimated ± 3.4 pp, medium confidence
97Gemini 3.5 Flash51.3%estimated ± 3.6 pp, medium confidence
98Gemini 3 Pro51.3%estimated ± 3.6 pp, medium confidence
99GLM-5.3-Flash51.2%measured
100GPT-5.6 Terra51.0%measured
101Apodex 1.150.8%estimated ± 3.4 pp, medium confidence
102Ornith-1.5-397B50.8%estimated ± 3.4 pp, medium confidence
103Qwen3.8 Max50.7%estimated ± 3.4 pp, medium confidence
104LFM2.5-8B-A1B50.4%estimated ± 3.6 pp, medium confidence
105Qwen3.8-27B50.0%estimated ± 5.7 pp, low confidence
106Claude Opus 4.849.9%estimated ± 3.6 pp, medium confidence
107Claude Haiku 5.549.6%estimated ± 3.4 pp, medium confidence
108Claude Sonnet 549.6%estimated ± 3.4 pp, medium confidence
109Qwen3.5-35B-A3B49.5%estimated ± 3.6 pp, medium confidence
110Claude Fable 5.149.5%measured
111GPT-6 Sol49.4%measured
112Claude Opus 4.5 Thinking49.3%estimated ± 3.6 pp, medium confidence
113Trinity-Large-Preview48.8%estimated ± 3.6 pp, medium confidence
114Trinity-Large-Thinking48.8%estimated ± 3.6 pp, medium confidence
115GPT-6 Astra48.6%measured
116MiMo-V2-Omni48.0%estimated ± 3.6 pp, medium confidence
117Muse Spark47.8%estimated ± 3.6 pp, medium confidence
118Kimi K347.7%measured
119Claude Opus 4.6 (Adaptive)47.4%estimated ± 3.6 pp, medium confidence
120GPT-5.2-Codex47.4%estimated ± 3.6 pp, medium confidence
121DeepSeek V4 Pro 081347.4%estimated ± 3.4 pp, medium confidence
122Inkling-Small47.2%estimated ± 3.6 pp, medium confidence
123DeepSeek V4.1 Flash46.9%measured
124Claude Opus 4.7 (Adaptive)46.7%measured
125Grok 4.1 Fast (Reasoning)46.6%estimated ± 3.6 pp, medium confidence
126Qwen3.5-122B-A10B46.4%estimated ± 3.6 pp, medium confidence
127Beam46.4%estimated ± 3.6 pp, medium confidence
128MiMo-V2.6-Pro46.3%estimated ± 3.9 pp, medium confidence
129Claude Mythos Preview46.3%estimated ± 3.9 pp, medium confidence
130GPT-5.6 Luna46.3%estimated ± 3.9 pp, low confidence
131GPT-6 Luna46.3%estimated ± 3.9 pp, low confidence
132MiMo-V2.6-Flash46.3%estimated ± 3.9 pp, low confidence
133Qwen3.5-27B46.2%estimated ± 3.6 pp, medium confidence
134GLM-5.346.1%measured
135MiMo-V2.5-Pro46.1%estimated ± 3.6 pp, medium confidence
136Mistral Medium 3.5 128B46.1%estimated ± 3.6 pp, medium confidence
137Qwen3.6-27B46.1%estimated ± 3.6 pp, medium confidence
138GPT-5.545.8%measured
139MiMo-V2-Pro45.6%estimated ± 3.6 pp, medium confidence
140Gemini 3.1 Pro45.2%estimated ± 3.6 pp, medium confidence
141GLM-4.745.1%estimated ± 3.6 pp, medium confidence
142Kimi K2.5 (Reasoning)45.1%estimated ± 3.6 pp, medium confidence
143Qwen 3.6 Max (preview)45.1%estimated ± 3.6 pp, medium confidence
144Grok 4.344.1%estimated ± 3.6 pp, medium confidence
145Gemini 3.7 Flash44.0%estimated ± 6.6 pp, low confidence
146Kimi K2.7 Code43.8%estimated ± 3.6 pp, medium confidence
147Claude Fable 543.7%estimated ± 3.6 pp, medium confidence
148GLM-5-Turbo43.7%estimated ± 3.6 pp, medium confidence
149GLM-5V-Turbo43.7%estimated ± 3.6 pp, medium confidence
150Claude Sonnet 5.543.4%estimated ± 3.4 pp, medium confidence
151Muse Glimmer 30B43.3%estimated ± 3.6 pp, medium confidence
152Claude Opus 543.2%estimated ± 3.4 pp, medium confidence
153GLM-5.242.7%measured
154MiniMax M342.1%estimated ± 3.6 pp, low confidence
155Grok 4.742.1%measured
156Inkling42.1%estimated ± 3.6 pp, low confidence
157Qwen3.7 Plus41.3%estimated ± 3.6 pp, low confidence
158GLM-5.140.0%estimated ± 3.6 pp, low confidence
159GPT-5.439.0%estimated ± 3.6 pp, low confidence
160Claude Opus 5.538.2%measured
161Ling 3.0 Flash34.8%estimated ± 3.6 pp, low confidence
162Grok 4.633.3%estimated ± 5.7 pp, low confidence
163Muse Spark 1.333.2%measured
164Qwen3.6-35B-A3B32.7%estimated ± 3.6 pp, low confidence
165Solar Pro 431.7%estimated ± 3.6 pp, low confidence
166Solar Open 229.3%estimated ± 3.6 pp, low confidence
167GPT-5.4 mini29.0%estimated ± 3.6 pp, low confidence
168GPT-5.4 nano27.9%estimated ± 3.6 pp, low confidence
169Kimi K2.627.7%estimated ± 3.6 pp, low confidence
170Qwen3.6 Plus22.7%estimated ± 3.6 pp, low confidence
171LLaDA2.2-flash21.5%estimated ± 3.6 pp, low confidence
172Qwen3.5 397B21.4%estimated ± 3.6 pp, low confidence
173LongCat-Flash-Lite-Sparse21.1%estimated ± 3.6 pp, low confidence
174Claude Opus 4.519.2%estimated ± 3.6 pp, low confidence
175GLM-513.2%estimated ± 3.6 pp, low confidence
176Kimi K2.512.4%estimated ± 3.6 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General