benchgap
Coding

PostTrainBench v1.1 leaderboard

As of 2026-10-07, the highest measured score on PostTrainBench v1.1 is 49.3% by Claude Opus 5.5. 175 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 5.549.3%measured
2Gemini 4 Argon45.3%measured
3Claude Sonnet 5.544.4%estimated ± 4.3 pp, medium confidence
4GPT-6 Astra44.3%measured
5Muse Spark 1.343.9%estimated ± 4.8 pp, high confidence
6Ember-143.7%estimated ± 4.8 pp, high confidence
7DeepSeek V4.1 Flash42.8%estimated ± 4.8 pp, high confidence
8Pareto 26.942.6%estimated ± 4.8 pp, high confidence
9SWE-241.6%estimated ± 4.8 pp, high confidence
10GPT-6.1 Sol40.6%estimated ± 4.8 pp, high confidence
11MiMo-V2.6-Pro40.6%estimated ± 4.8 pp, high confidence
12Claude Fable 5.140.2%measured
13Claude Mythos 540.1%estimated ± 4.7 pp, high confidence
14Pareto 26.10 Preview38.7%estimated ± 4.8 pp, high confidence
15GPT-6 Sol37.6%estimated ± 4.8 pp, high confidence
16MiMo-V2.6-Flash36.7%estimated ± 4.8 pp, high confidence
17Step 5 Preview36.5%estimated ± 4.8 pp, high confidence
18GPT-5.6 Sol36.2%measured
19GLM-5.336.2%estimated ± 2.7 pp, high confidence
20GPT-6 Luna35.5%estimated ± 4.8 pp, high confidence
21Claude Opus 535.0%measured
22Sakana Fugu-Ultra34.9%estimated ± 4.7 pp, high confidence
23DeepSeek V4 Pro 081334.5%estimated ± 2.7 pp, high confidence
24GPT-5.6 Terra34.4%estimated ± 2.7 pp, high confidence
25GLM-5.3-Flash34.1%estimated ± 2.7 pp, high confidence
26Grok 4.633.8%estimated ± 2.7 pp, high confidence
27Grok 4.733.4%estimated ± 4.3 pp, medium confidence
28GPT-5.6 Luna33.1%estimated ± 3.3 pp, high confidence
29Claude Fable 532.9%estimated ± 2.7 pp, high confidence
30Claude Opus 4.832.9%measured
31Claude Haiku 5.532.6%estimated ± 5.9 pp, low confidence
32Kimi K332.0%measured
33GLM-5.231.7%measured
34Claude Haiku 4.530.9%estimated ± 2.7 pp, high confidence
35Mistral Large 430.4%estimated ± 7.2 pp, low confidence
36Ling 3.1 Flash30.3%estimated ± 7.2 pp, low confidence
37DeepSeek V4 Flash 073129.8%estimated ± 2.7 pp, high confidence
38Hy4 preview29.3%estimated ± 4.7 pp, high confidence
39Qwen3.8-27B29.1%estimated ± 2.7 pp, high confidence
40Beam29.1%estimated ± 4.7 pp, high confidence
41Muse Spark 1.229.1%estimated ± 3.3 pp, high confidence
42Ornith-1.5-397B28.9%estimated ± 4.7 pp, high confidence
43SWE-1.728.8%estimated ± 5.9 pp, low confidence
44Claude Opus 4.728.6%measured
45Claude Opus 4.7 (Adaptive)28.4%estimated ± 4.7 pp, high confidence
46Qwen3.8-Omni-Flash27.7%estimated ± 4.7 pp, high confidence
47Qwen3.8 Max27.6%estimated ± 2.7 pp, high confidence
48GPT-5.527.2%measured
49Qwen3.8-Flash-Next27.2%estimated ± 4.7 pp, high confidence
50Qwen3.8 Max Preview27.1%estimated ± 6.1 pp, medium confidence
51Ornith-1.0-397B27.0%estimated ± 4.7 pp, high confidence
52dots3-note Preview26.3%estimated ± 4.7 pp, high confidence
53Inkling-Small26.1%estimated ± 2.7 pp, high confidence
54Muse Spark 1.126.0%estimated ± 2.7 pp, high confidence
55Claude Sonnet 525.6%estimated ± 2.7 pp, high confidence
56Atria Dawn Preview25.4%estimated ± 4.7 pp, high confidence
57Ornith-1.5-35B-A3B25.4%estimated ± 4.7 pp, high confidence
58Laguna S 2.125.3%estimated ± 4.7 pp, high confidence
59Composer 2.525.2%estimated ± 3.3 pp, high confidence
60Sakana Fugu25.0%estimated ± 4.7 pp, high confidence
61Kimi K2.7 Code24.9%estimated ± 2.7 pp, high confidence
62Claude Sonnet 4.624.4%estimated ± 2.7 pp, high confidence
63Solar Pro 424.4%estimated ± 7.2 pp, low confidence
64Gemini 3.7 Flash24.3%estimated ± 2.7 pp, high confidence
65Ling 3.0 Flash VL24.2%estimated ± 7.2 pp, low confidence
66Qwen3.7 Plus24.2%estimated ± 4.7 pp, medium confidence
67GLM-5.124.0%estimated ± 2.7 pp, high confidence
68Claude Opus 4.523.9%estimated ± 4.7 pp, medium confidence
69Gemini 3.5 Flash-Lite23.9%estimated ± 2.7 pp, high confidence
70Gemini 3.6 Flash23.8%estimated ± 2.7 pp, high confidence
71Gemini 3.8 Flash23.6%estimated ± 2.7 pp, high confidence
72K-EXAONE 2.023.5%estimated ± 7.2 pp, low confidence
73Grok 4.523.5%measured
74Gemini 3.5 Flash23.4%estimated ± 2.7 pp, high confidence
75Step 3.7 Flash23.4%estimated ± 4.7 pp, medium confidence
76Inkling23.4%estimated ± 2.7 pp, high confidence
77A.X K223.2%estimated ± 7.2 pp, low confidence
78GPT-5.3 Codex23.0%estimated ± 2.7 pp, high confidence
79GPT-5.223.0%estimated ± 4.7 pp, medium confidence
80MiniMax M322.8%estimated ± 2.7 pp, high confidence
81MiniMax M2.722.8%estimated ± 2.7 pp, high confidence
82GLM-522.8%estimated ± 4.7 pp, medium confidence
83DeepSeek V3 032422.7%estimated ± 7.2 pp, low confidence
84North Mini Code22.7%estimated ± 7.2 pp, low confidence
85Mercury 2.522.6%estimated ± 7.2 pp, low confidence
86Muse Spark22.6%estimated ± 3.3 pp, medium confidence
87MiMo-V2.5-Pro22.5%estimated ± 2.7 pp, high confidence
88Kimi K2.622.1%estimated ± 2.7 pp, high confidence
89Gemini 3.1 Pro22.0%measured
90Claude Opus 4.621.8%estimated ± 4.7 pp, medium confidence
91Qwen 3.6 Max (preview)21.8%estimated ± 3.3 pp, medium confidence
92GPT-5.4 mini21.8%estimated ± 2.7 pp, high confidence
93Gemini 3 Flash21.7%estimated ± 2.7 pp, high confidence
94Solar Pro 321.5%estimated ± 7.2 pp, low confidence
95MAI-Thinking-121.5%estimated ± 4.7 pp, medium confidence
96Granite 4.2 3B21.5%estimated ± 7.2 pp, low confidence
97Ling 3.0 Tiny21.4%estimated ± 7.2 pp, low confidence
98Claude Opus 4.6 (Adaptive)21.1%estimated ± 10.5 pp, low confidence
99Muse Glimmer 30B20.6%estimated ± 4.7 pp, medium confidence
100Qwen3.6 Plus20.6%estimated ± 2.7 pp, high confidence
101MiMo-V2.520.5%estimated ± 2.7 pp, high confidence
102Qwen3.6-27B20.5%estimated ± 3.3 pp, medium confidence
103Qwen3.5 397B20.4%estimated ± 4.7 pp, medium confidence
104Grok 4.2020.4%estimated ± 2.7 pp, high confidence
105Kimi K2.520.3%estimated ± 4.7 pp, medium confidence
106Ornith-1.0-35B20.2%estimated ± 4.7 pp, medium confidence
107GLM-5 (Reasoning)19.9%estimated ± 10.5 pp, low confidence
108Claude Sonnet 4.5 Thinking19.9%estimated ± 10.5 pp, low confidence
109GPT-5.1-Codex-Max19.9%estimated ± 10.5 pp, low confidence
110Claude Opus 4.5 Thinking19.9%estimated ± 10.5 pp, low confidence
111Qwen3.5 Plus19.9%estimated ± 10.5 pp, low confidence
112MiniMax M2.519.9%estimated ± 10.5 pp, low confidence
113Gemini 3 Pro19.9%estimated ± 10.5 pp, low confidence
114GPT-5 mini19.9%estimated ± 10.5 pp, low confidence
115GPT-5.1-Codex19.9%estimated ± 10.5 pp, low confidence
116Claude Haiku 4.5 Thinking19.9%estimated ± 10.5 pp, low confidence
117DeepSeek V3.2 (Thinking)19.9%estimated ± 10.5 pp, low confidence
118GLM-4.619.9%estimated ± 10.5 pp, low confidence
119Grok 4.1 Fast (Reasoning)19.9%estimated ± 10.5 pp, low confidence
120Grok 4 Fast (Reasoning)19.9%estimated ± 10.5 pp, low confidence
121Qwen3 Max19.9%estimated ± 10.5 pp, low confidence
122Grok 4.319.8%estimated ± 2.7 pp, high confidence
123Qwen3.6-35B-A3B19.7%estimated ± 4.7 pp, medium confidence
124GLM-4.719.3%estimated ± 2.7 pp, high confidence
125GPT-5.2-Codex19.3%estimated ± 2.7 pp, high confidence
126GPT-5.419.0%measured
127GPT-5.4 nano19.0%estimated ± 2.7 pp, high confidence
128Mistral Medium 3.5 128B18.9%estimated ± 3.3 pp, medium confidence
129Laguna XS 2.118.7%estimated ± 4.7 pp, medium confidence
130Ornith-1.5-9B18.6%estimated ± 4.7 pp, medium confidence
131Nemotron 3 Ultra17.8%estimated ± 2.7 pp, high confidence
132Qwen3.7 Max17.3%estimated ± 2.7 pp, high confidence
133Laguna M.116.5%estimated ± 2.7 pp, high confidence
134Quasar 438B16.3%estimated ± 6.1 pp, low confidence
135Ornith-1.0-9B16.3%estimated ± 4.7 pp, medium confidence
136Ling 3.0 Flash16.1%estimated ± 2.7 pp, high confidence
137Apodex 1.116.0%estimated ± 6.1 pp, low confidence
138Apodex 1.1 Mini16.0%estimated ± 6.1 pp, low confidence
139Gemini 3.1 Flash-Lite15.8%estimated ± 2.7 pp, high confidence
140Qwen3.5 Flash15.8%estimated ± 2.7 pp, high confidence
141LongCat-Flash-Lite-Sparse15.3%estimated ± 4.7 pp, medium confidence
142Laguna XS.215.1%estimated ± 2.7 pp, high confidence
143Hy314.7%estimated ± 6.1 pp, low confidence
144Hy3 Preview14.7%estimated ± 6.1 pp, low confidence
145Gemini 2.5 Pro14.2%estimated ± 3.3 pp, medium confidence
146Granite 4.2 30B11.9%estimated ± 4.7 pp, medium confidence
147Ling 3.0 Flash FP811.0%estimated ± 6.1 pp, low confidence
148MiMo-V2-Flash10.8%estimated ± 6.1 pp, low confidence
149GPT-5.110.7%estimated ± 6.1 pp, low confidence
150LLaDA2.2-flash10.6%estimated ± 4.7 pp, medium confidence
151Kimi K2.5 (Reasoning)10.0%estimated ± 6.1 pp, low confidence
152Qwen3.5-122B-A10B9.8%estimated ± 6.1 pp, low confidence
153Gemma 4 31B9.4%estimated ± 6.1 pp, low confidence
154o19.0%estimated ± 6.1 pp, low confidence
155Gemma 4 26B A4B8.9%estimated ± 6.1 pp, low confidence
156GPT-5 (high)8.8%estimated ± 6.1 pp, low confidence
157Nemotron 3 Super 100B8.8%estimated ± 6.1 pp, low confidence
158o1-preview8.6%estimated ± 6.1 pp, low confidence
159K-Exaone8.5%estimated ± 6.1 pp, low confidence
160Gemma 4 12B8.5%estimated ± 6.1 pp, low confidence
161GPT-OSS 120B8.4%estimated ± 6.1 pp, low confidence
162Command A+8.4%estimated ± 6.1 pp, low confidence
163Nemotron 3.5 Lightning 30B A3B NVFP48.4%estimated ± 6.1 pp, low confidence
164Mistral Small 48.4%estimated ± 6.1 pp, low confidence
165Mistral Small 4 (Reasoning)8.4%estimated ± 6.1 pp, low confidence
166Trinity-Large-Preview8.4%estimated ± 6.1 pp, low confidence
167Trinity-Large-Thinking8.4%estimated ± 6.1 pp, low confidence
168Ling 2.6 Flash8.4%estimated ± 6.1 pp, low confidence
169Gemini 1.5 Pro8.3%estimated ± 6.1 pp, low confidence
170DeepSeek V38.3%estimated ± 6.1 pp, low confidence
171GPT-4 Turbo8.3%estimated ± 6.1 pp, low confidence
172GPT-OSS 20B8.3%estimated ± 6.1 pp, low confidence
173GPT-4.1 mini8.3%estimated ± 6.1 pp, low confidence
174Mistral Large 38.3%estimated ± 6.1 pp, low confidence
175Claude 3 Opus8.3%estimated ± 6.1 pp, low confidence
176Llama 4 Maverick8.3%estimated ± 6.1 pp, low confidence
177Celeris-18.3%estimated ± 6.1 pp, low confidence
178Nemotron 3 Nano 30B8.3%estimated ± 6.1 pp, low confidence
179Nemotron 3 Nano Omni 30B A3B8.3%estimated ± 6.1 pp, low confidence
180Ultravox v0.6 Llama 3.3 70B8.3%estimated ± 6.1 pp, low confidence
181GPT-4o mini8.3%estimated ± 6.1 pp, low confidence
182GPT-4.1 nano8.3%estimated ± 6.1 pp, low confidence
183Gemma 3 27B8.3%estimated ± 6.1 pp, low confidence
184Gemma 4 E4B8.3%estimated ± 6.1 pp, low confidence
185Gemma 4 E2B8.3%estimated ± 6.1 pp, low confidence
186LFM2.5-2.6B8.3%estimated ± 6.1 pp, low confidence
187Llama 4 Scout8.3%estimated ± 6.1 pp, low confidence
188Granite 4.2 8B6.3%estimated ± 4.7 pp, medium confidence
189MiniCPM5-2B4.6%estimated ± 4.7 pp, medium confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General