benchgap
Knowledge & reasoning

SimpleQA leaderboard

As of 2026-10-10, the highest measured score on SimpleQA is 54.3% by Qwen3 235B 2507. 101 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1GPT-6 Astra76.9%estimated ± 10.9 pp, low confidence
2Sakana Fugu76.1%estimated ± 10.9 pp, low confidence
3Sakana Fugu-Ultra76.1%estimated ± 10.9 pp, low confidence
4GPT-5.6 Sol74.7%estimated ± 10.9 pp, low confidence
5Claude Opus 4.7 (Adaptive)74.0%estimated ± 10.9 pp, low confidence
6Claude Mythos 573.8%estimated ± 10.9 pp, low confidence
7Claude Opus 4.873.0%estimated ± 10.9 pp, low confidence
8GPT-5.573.0%estimated ± 10.9 pp, low confidence
9Kimi K372.8%estimated ± 10.9 pp, low confidence
10GPT-5.6 Terra71.8%estimated ± 10.9 pp, low confidence
11GPT-5.471.7%estimated ± 10.9 pp, low confidence
12Ornith-1.5-397B71.7%estimated ± 10.9 pp, low confidence
13Qwen3.8 Max71.3%estimated ± 10.9 pp, low confidence
14GPT-5.271.0%estimated ± 10.9 pp, low confidence
15Qwen3.7 Max71.0%estimated ± 10.9 pp, low confidence
16GPT-5.6 Luna70.8%estimated ± 10.9 pp, low confidence
17Hy4 preview70.8%estimated ± 10.9 pp, low confidence
18Gemini 3.5 Flash70.6%estimated ± 10.9 pp, low confidence
19Qwen3.8-Flash-Next69.8%estimated ± 10.9 pp, low confidence
20Claude Opus 4.669.1%estimated ± 10.9 pp, low confidence
21GLM-5.268.9%estimated ± 10.9 pp, low confidence
22Qwen3.8-Omni-Flash68.6%estimated ± 10.9 pp, low confidence
23DeepSeek V4.1 Flash68.4%estimated ± 10.9 pp, low confidence
24Kimi K2.667.7%estimated ± 10.9 pp, low confidence
25Qwen3.6 Plus67.5%estimated ± 10.9 pp, low confidence
26Qwen3.7 Plus67.3%estimated ± 10.9 pp, low confidence
27DeepSeek V4 Pro 081367.0%estimated ± 10.9 pp, low confidence
28Grok 4.367.0%estimated ± 10.9 pp, low confidence
29Claude Sonnet 4.666.6%estimated ± 10.9 pp, low confidence
30Interfaze Beta66.6%estimated ± 10.9 pp, low confidence
31Inkling-Small65.9%estimated ± 10.9 pp, low confidence
32Ornith-1.5-35B-A3B65.4%estimated ± 10.9 pp, low confidence
33Qwen3.8-27B65.4%estimated ± 10.9 pp, low confidence
34Qwen3.5 397B64.0%estimated ± 10.9 pp, low confidence
35DeepSeek V4 Flash 073163.4%estimated ± 10.9 pp, low confidence
36GPT-5.4 mini63.2%estimated ± 10.9 pp, low confidence
37Inkling63.1%estimated ± 10.9 pp, low confidence
38Qwen3.6-27B62.9%estimated ± 10.9 pp, low confidence
39Kimi K2.562.5%estimated ± 10.9 pp, low confidence
40Kimi K2.5 (Reasoning)62.5%estimated ± 10.9 pp, low confidence
41Hy3 Preview61.8%estimated ± 10.9 pp, low confidence
42Claude Opus 4.561.4%estimated ± 10.9 pp, low confidence
43Nemotron 3 Ultra61.4%estimated ± 10.9 pp, low confidence
44Qwen3.5-122B-A10B60.7%estimated ± 10.9 pp, low confidence
45Ornith-1.5-9B60.4%estimated ± 10.9 pp, low confidence
46GLM-559.6%estimated ± 10.9 pp, low confidence
47Qwen3.6-35B-A3B59.6%estimated ± 10.9 pp, low confidence
48Ternary Bonsai 2 27B59.2%estimated ± 10.9 pp, low confidence
49GLM-4.759.1%estimated ± 10.9 pp, low confidence
50Qwen3.5-27B58.7%estimated ± 10.9 pp, low confidence
51Ling 3.0 Flash57.8%estimated ± 10.9 pp, low confidence
52Gemma 4 31B56.6%estimated ± 10.9 pp, low confidence
53MAI-Thinking-156.4%estimated ± 10.9 pp, low confidence
54Qwen3.5-35B-A3B56.4%estimated ± 10.9 pp, low confidence
55Ling 3.0 Flash FP856.0%estimated ± 10.9 pp, low confidence
56MiMo-V2-Flash55.5%estimated ± 10.9 pp, low confidence
57Claude Sonnet 4.554.9%estimated ± 10.9 pp, low confidence
58Qwen3 235B 250754.3%measured
59Gemini 2.5 Pro54.2%estimated ± 10.9 pp, low confidence
60GPT-5.4 nano53.9%estimated ± 10.9 pp, low confidence
61Gemma 4 26B A4B53.0%estimated ± 10.9 pp, low confidence
62Qwen3 235B 2507 (Reasoning)50.8%estimated ± 10.9 pp, low confidence
63Grok 3 Mini49.4%estimated ± 10.9 pp, low confidence
64DeepSeek V3.1 (Reasoning)49.1%estimated ± 10.9 pp, low confidence
65GPT-OSS 120B49.1%estimated ± 10.9 pp, low confidence
66Nemotron 3 Super 100B47.6%estimated ± 10.9 pp, low confidence
67o1-pro47.2%estimated ± 10.9 pp, low confidence
68Gemma 4 12B46.8%estimated ± 10.9 pp, low confidence
69o3-mini44.1%estimated ± 10.9 pp, low confidence
70Grok 3 [Beta]43.6%measured
71Nemotron Ultra 253B42.2%estimated ± 10.9 pp, low confidence
72o141.6%estimated ± 10.9 pp, low confidence
73Nemotron 3.5 Lightning 30B A3B NVFP441.4%estimated ± 10.9 pp, low confidence
74GLM-4.7-Flash40.8%estimated ± 10.9 pp, low confidence
75DeepSeek V3.140.4%estimated ± 10.9 pp, low confidence
76Nemotron 3 Nano 30B37.5%estimated ± 10.9 pp, low confidence
77Nemotron 3 Nano Omni 30B A3B36.3%estimated ± 10.9 pp, low confidence
78GPT-OSS 20B35.3%estimated ± 10.9 pp, low confidence
79Ministral 3 14B (Reasoning)34.9%estimated ± 10.9 pp, low confidence
80ZAYA1-8B34.6%estimated ± 10.9 pp, low confidence
81Kimi K231.0%measured
82DeepSeek-R130.1%measured
83Granite 4.2 30B28.8%estimated ± 10.9 pp, low confidence
84GPT-4.128.7%estimated ± 10.9 pp, low confidence
85GPT-4.1 mini26.4%estimated ± 10.9 pp, low confidence
86Granite 4.2 8B26.4%estimated ± 10.9 pp, low confidence
87DeepSeek V324.9%measured
88Claude 3.5 Sonnet22.2%estimated ± 10.9 pp, low confidence
89Ling 2.6 Flash21.9%estimated ± 10.9 pp, low confidence
90Gemma 4 E4B21.6%estimated ± 10.9 pp, low confidence
91Mellum2-12B-A2.5B-Thinking21.0%estimated ± 10.9 pp, low confidence
92ZAYA1-74B-Preview20.8%estimated ± 10.9 pp, low confidence
93Phi-420.0%estimated ± 10.9 pp, low confidence
94Granite 4.2 3B19.3%estimated ± 10.9 pp, low confidence
95GPT-4.1 nano17.3%estimated ± 10.9 pp, low confidence
96Qwen2.5-VL-32B16.1%estimated ± 10.9 pp, low confidence
97Gemma 4 E2B15.5%estimated ± 10.9 pp, low confidence
98Soofi S 30B-A3B15.5%estimated ± 10.9 pp, low confidence
99Mellum2-12B-A2.5B-Instruct15.1%estimated ± 10.9 pp, low confidence
100LFM2.5-350M14.4%estimated ± 10.9 pp, low confidence
101Granite-4.0-1B14.4%estimated ± 10.9 pp, low confidence
102Granite-4.0-H-1B14.4%estimated ± 10.9 pp, low confidence
103Granite-4.0-H-350M14.3%estimated ± 10.9 pp, low confidence
104LFM2.5-VL-450M14.3%estimated ± 10.9 pp, low confidence
105LFM2.5-230M14.3%estimated ± 10.9 pp, low confidence
106Granite-4.0-350M14.3%estimated ± 10.9 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General