benchgap
Agentic · tools

AA Harvey LAB leaderboard

As of 2026-10-07, the highest measured score on AA Harvey LAB is 94.6% by Kimi K3. 99 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Kimi K394.6%measured
2Muse Spark 1.393.7%estimated ± 1.2 pp, medium confidence
3GLM-5.393.6%estimated ± 1.2 pp, medium confidence
4Grok 4.693.6%estimated ± 1.2 pp, medium confidence
5Claude Fable 593.6%measured
6GPT-6 Astra93.5%estimated ± 1.2 pp, medium confidence
7Atria Dawn Preview93.5%estimated ± 9.8 pp, low confidence
8Claude Opus 593.5%measured
9GPT-5.6 Sol93.5%estimated ± 1.2 pp, medium confidence
10Qwen3.8 Max Preview93.4%estimated ± 1.2 pp, medium confidence
11DeepSeek V4 Pro 081393.4%estimated ± 1.2 pp, medium confidence
12Step 5 Preview93.4%measured
13Gemini 4 Argon93.3%estimated ± 1.3 pp, medium confidence
14Qwen3.8-27B93.2%estimated ± 1.2 pp, medium confidence
15Claude Sonnet 5.593.1%measured
16Claude Fable 5.193.0%measured
17GPT-5.5 Pro92.9%estimated ± 9.8 pp, low confidence
18GPT-6.1 Sol92.9%estimated ± 2.1 pp, high confidence
19Claude Sonnet 592.9%estimated ± 1.2 pp, medium confidence
20Muse Spark 1.292.9%estimated ± 1.2 pp, medium confidence
21GPT-5.6 Terra92.9%estimated ± 1.2 pp, medium confidence
22GPT-5.4 Pro92.8%estimated ± 9.8 pp, low confidence
23GPT-5.6 Luna92.7%estimated ± 1.2 pp, medium confidence
24Claude Opus 4.892.7%estimated ± 1.2 pp, medium confidence
25GPT-6 Sol92.7%estimated ± 1.3 pp, medium confidence
26Grok 4.592.6%estimated ± 1.2 pp, medium confidence
27DeepSeek V4 Flash 073192.5%estimated ± 1.2 pp, medium confidence
28Claude Mythos 592.4%estimated ± 9.8 pp, low confidence
29Gemini 3.8 Flash92.4%estimated ± 1.2 pp, medium confidence
30Claude Opus 4.7 (Adaptive)92.1%estimated ± 1.2 pp, medium confidence
31GLM-5.292.0%estimated ± 1.2 pp, medium confidence
32DeepSeek V4.1 Flash92.0%estimated ± 1.7 pp, medium confidence
33GPT-5.591.4%estimated ± 1.2 pp, medium confidence
34Claude Opus 5.591.2%measured
35Gemini 3.7 Flash90.7%measured
36Claude Haiku 5.589.9%measured
37Beam89.9%estimated ± 9.8 pp, low confidence
38Apodex 1.189.8%estimated ± 1.7 pp, medium confidence
39Ornith-1.5-397B89.8%estimated ± 1.7 pp, medium confidence
40Qwen3.8-Flash-Next89.7%estimated ± 1.3 pp, medium confidence
41Qwen3.8 Max89.7%estimated ± 1.3 pp, medium confidence
42Hy4 preview89.6%estimated ± 1.7 pp, medium confidence
43Quasar 438B89.4%estimated ± 1.2 pp, medium confidence
44Claude Opus 4.789.3%estimated ± 1.3 pp, medium confidence
45Claude Sonnet 4.688.9%estimated ± 1.3 pp, medium confidence
46GLM-5.3-Flash88.8%estimated ± 1.3 pp, medium confidence
47dots3-note Preview88.7%estimated ± 1.7 pp, medium confidence
48MiniMax M388.4%measured
49Gemini 3.6 Flash87.7%estimated ± 1.2 pp, medium confidence
50Agents-A1-4B87.2%estimated ± 9.8 pp, low confidence
51GPT-5.286.9%estimated ± 9.8 pp, low confidence
52Agents-A186.8%estimated ± 1.7 pp, medium confidence
53Step 3.7 Flash86.6%estimated ± 1.7 pp, medium confidence
54Qwen3.5 397B86.0%estimated ± 9.8 pp, low confidence
55Muse Spark 1.185.8%estimated ± 1.2 pp, medium confidence
56Qwen3.5-27B85.7%estimated ± 9.8 pp, low confidence
57Qwen3.5-35B-A3B85.7%estimated ± 9.8 pp, low confidence
58Kimi K2.585.6%estimated ± 9.8 pp, low confidence
59Kimi K2.5 (Reasoning)85.6%estimated ± 9.8 pp, low confidence
60Gemini 3.5 Flash85.6%estimated ± 1.2 pp, medium confidence
61Hy384.3%estimated ± 1.2 pp, medium confidence
62Hy3 Preview84.3%estimated ± 1.2 pp, medium confidence
63GLM-5.184.0%estimated ± 1.2 pp, medium confidence
64Inkling-Small83.8%estimated ± 1.2 pp, medium confidence
65Inkling83.4%estimated ± 1.2 pp, medium confidence
66GLM-4.783.3%estimated ± 9.8 pp, low confidence
67Qwen3.7 Max83.1%estimated ± 1.2 pp, medium confidence
68Solar Pro 482.6%estimated ± 9.8 pp, low confidence
69LongCat-Flash-Lite-Sparse82.4%estimated ± 9.8 pp, low confidence
70MiMo-V2.5-Pro82.3%estimated ± 1.2 pp, medium confidence
71Kimi K2.7 Code82.2%estimated ± 1.2 pp, medium confidence
72Kimi K2.682.0%estimated ± 1.2 pp, medium confidence
73Nemotron 3 Ultra81.7%measured
74Ling 3.0 Flash81.4%estimated ± 1.2 pp, low confidence
75Ling 3.0 Flash FP881.4%estimated ± 1.2 pp, low confidence
76Qwen3.6-27B81.0%estimated ± 1.2 pp, low confidence
77Qwen3.7 Plus80.8%estimated ± 1.2 pp, low confidence
78GPT-5.4 mini80.8%estimated ± 1.2 pp, low confidence
79GPT-5.4 nano80.2%estimated ± 1.2 pp, low confidence
80Grok 4.380.1%estimated ± 1.2 pp, low confidence
81MiniMax M2.780.0%estimated ± 1.2 pp, low confidence
82Gemini 3.5 Flash-Lite79.8%estimated ± 1.2 pp, low confidence
83Qwen3.6-35B-A3B79.7%estimated ± 1.2 pp, low confidence
84Muse Glimmer 30B79.4%estimated ± 1.2 pp, low confidence
85Gemini 3.1 Pro79.4%estimated ± 1.2 pp, low confidence
86Qwen3.5-122B-A10B79.4%estimated ± 1.2 pp, low confidence
87Mistral Medium 3.5 128B79.4%estimated ± 1.2 pp, low confidence
88Gemma 4 31B79.4%estimated ± 1.2 pp, low confidence
89GPT-OSS 120B79.4%estimated ± 1.2 pp, low confidence
90Nemotron 3.5 Lightning 30B A3B NVFP479.4%estimated ± 1.2 pp, low confidence
91Nemotron 3 Super 100B79.4%estimated ± 1.2 pp, low confidence
92Command A+79.4%estimated ± 1.2 pp, low confidence
93Gemini 2.5 Pro79.4%estimated ± 1.2 pp, low confidence
94Granite 4.2 8B79.4%estimated ± 1.2 pp, low confidence
95Celeris-179.4%estimated ± 1.2 pp, low confidence
96DeepSeek V379.4%estimated ± 1.2 pp, low confidence
97Gemma 3 27B79.4%estimated ± 1.2 pp, low confidence
98GPT-OSS 20B79.4%estimated ± 1.2 pp, low confidence
99Llama 4 Maverick79.4%estimated ± 1.2 pp, low confidence
100Llama 4 Scout79.4%estimated ± 1.2 pp, low confidence
101Mistral Large 379.4%estimated ± 1.2 pp, low confidence
102Mistral Small 479.4%estimated ± 1.2 pp, low confidence
103Mistral Small 4 (Reasoning)79.4%estimated ± 1.2 pp, low confidence
104Nemotron 3 Nano 30B79.4%estimated ± 1.2 pp, low confidence
105Trinity-Large-Preview79.4%estimated ± 1.2 pp, low confidence
106Trinity-Large-Thinking79.4%estimated ± 1.2 pp, low confidence
107Ornith-1.5-35B-A3B79.2%estimated ± 1.7 pp, low confidence
108Ornith-1.5-9B77.1%estimated ± 1.7 pp, low confidence
109Grok 4.719.6%measured
110GPT-5.40.1%estimated ± 2.0 pp, low confidence
111Claude Opus 4.60.0%estimated ± 2.0 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General