benchgap
Agentic · tools

Harvey LAB leaderboard

As of 2026-10-07, the highest measured score on Harvey LAB is 95.5% by Muse Spark 1.3 (xhigh). 34 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: artificial-analysis.

#ModelScoreSource
1Gemini 4 Argon (high)95.6%estimated ± 3.2 pp, medium confidence
2Muse Spark 1.3 (xhigh)95.5%measured
3Kimi K3 (max)94.6%measured
4K2 Horizon 375B A23B94.2%estimated ± 2.1 pp, medium confidence
5Qwen3.8 Max94.2%estimated ± 2.1 pp, low confidence
6Qwen3.8 2.4T A95B94.2%estimated ± 2.1 pp, low confidence
7Qwen3.8 27B (medium)94.2%estimated ± 2.1 pp, medium confidence
8Gemini 3.7 Flash (medium)94.1%estimated ± 1.2 pp, medium confidence
9Gemini 3.5 Flash (high)94.1%estimated ± 1.2 pp, medium confidence
10DeepSeek V4 Pro 0813 (max)94.0%estimated ± 1.2 pp, medium confidence
11GPT-6 Astra (xhigh)94.0%estimated ± 3.2 pp, high confidence
12Grok 4.6 (high)93.9%estimated ± 1.2 pp, medium confidence
13GPT-5.6 Sol (max)93.9%estimated ± 1.6 pp, low confidence
14Grok 4.6 (xhigh)93.9%estimated ± 3.2 pp, high confidence
15GPT-6 Astra (high)93.9%estimated ± 3.2 pp, high confidence
16GPT-6.1 Sol (xhigh)93.9%estimated ± 3.2 pp, high confidence
17Gemini 3.8 Flash (high)93.8%estimated ± 1.6 pp, low confidence
18GPT-5.6 Terra (max)93.8%estimated ± 1.6 pp, low confidence
19GPT-6 Sol (max)93.7%estimated ± 1.6 pp, low confidence
20Gemini 3.7 Flash (high)93.7%estimated ± 1.8 pp, medium confidence
21GPT-6 Astra (max)93.6%estimated ± 1.6 pp, low confidence
22Qwen3.8 Max (0902)93.6%measured
23Claude Fable 5 (with fallback)93.6%measured
24DeepSeek V4.1 Flash (max)93.5%estimated ± 1.6 pp, low confidence
25Claude Opus 4.7 (max)93.5%estimated ± 1.6 pp, low confidence
26Claude Opus 5 (max)93.5%measured
27Qwen3.8 27B (xhigh)93.5%estimated ± 1.2 pp, medium confidence
28GPT-5.5 (xhigh)93.4%estimated ± 1.6 pp, low confidence
29Step 5 Preview93.4%measured
30GPT-6.1 Sol (max)93.4%estimated ± 1.8 pp, high confidence
31Claude Fable 5.1 (xhigh with fallback)93.3%measured
32Claude Sonnet 5.5 (max with fallback)93.1%measured
33GLM-5.2 (max)93.0%estimated ± 1.6 pp, low confidence
34Claude Fable 5.1 (max with fallback)93.0%measured
35Claude Fable 5.1 (high with fallback)93.0%measured
36Grok 4.7 (xhigh)92.9%estimated ± 1.6 pp, low confidence
37Mistral Large 4 Preview92.8%estimated ± 3.2 pp, high confidence
38MiMo-V2.6-Pro92.6%estimated ± 3.2 pp, high confidence
39Inkling (xhigh)91.9%estimated ± 1.2 pp, medium confidence
40GPT-6 Luna (max)91.8%estimated ± 3.2 pp, high confidence
41Claude Opus 5.5 (max with fallback)91.2%measured
42GLM-5.3 (max)91.2%estimated ± 1.2 pp, medium confidence
43Muse Glimmer (high)90.1%estimated ± 1.2 pp, medium confidence
44Muse Spark 1.3 (max)89.3%estimated ± 1.6 pp, low confidence
45GLM-5.3-Flash88.8%estimated ± 1.2 pp, medium confidence
46MiniMax-M388.4%measured
47Nemotron 3 Ultra81.7%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Vision & documents