benchgap
Alibaba (Qwen) · model

Qwen3 Max benchmark scores

As of 2026-10-07, Qwen3 Max (Alibaba (Qwen)) has measured scores on 10 benchmarks and estimated scores on 53 more.

BenchmarkScoreSource
AA-GPQA Diamond76.4%measured
AA-HLE11.9%measured
AA-SciCode46.2%estimated ± 3.8 pp, high confidence
Artificial Analysis Intelligence Index15.6%measured
CritPt0.0%measured
GDPval-AA18.7%estimated ± 8.8 pp, medium confidence
GPQA Diamond75.6%estimated ± 2.2 pp, high confidence
HLE19.0%estimated ± 9.5 pp, medium confidence
IFBench56.3%estimated ± 6.9 pp, medium confidence
AA-LCR50.0%measured
AA-Omniscience Accuracy24.4%measured
BrowseComp59.8%estimated ± 5.7 pp, medium confidence
IFEval90.0%estimated ± 1.3 pp, high confidence
LongBench v20.0%estimated ± 5.9 pp, low confidence
τ²-bench results74.3%measured
AA Coding Index40.5%estimated ± 6.8 pp, medium confidence
APEX-Agents-AA1.6%estimated ± 11.4 pp, low confidence
BioMysteryBench (human-difficult)30.8%estimated ± 5.9 pp, low confidence
SWE-bench Verified72.3%estimated ± 4.4 pp, high confidence
BFCL v452.9%estimated ± 9.8 pp, low confidence
CyberGym36.9%estimated ± 10.0 pp, low confidence
τ³-bench results68.9%estimated ± 8.9 pp, medium confidence
WideResearch66.0%estimated ± 5.9 pp, low confidence
AA Agentic Index12.6%estimated ± 6.4 pp, medium confidence
MMLU-Pro85.5%estimated ± 3.5 pp, high confidence
GPQA79.7%estimated ± 3.2 pp, high confidence
AA-IFBench44.1%measured
Gert Labs43.7%measured
AA EnterpriseOps-Gym15.9%estimated ± 3.6 pp, low confidence
ApprenticeBench0.0%estimated ± 3.1 pp, low confidence
ARC-AGI-121.3%estimated ± 5.4 pp, low confidence
ARC-AGI-26.4%estimated ± 14.2 pp, low confidence
FrontierCode 1.1 Main1.3%estimated ± 4.3 pp, low confidence
OSWorld-Verified56.1%estimated ± 5.9 pp, medium confidence
Vals GPQA Diamond67.0%estimated ± 7.4 pp, medium confidence
Vals MMLU-Pro83.9%estimated ± 2.2 pp, high confidence
Vals SWE-bench64.1%estimated ± 6.3 pp, medium confidence
AA ITBench65.4%estimated ± 3.6 pp, low confidence
HLE w/o tools9.7%estimated ± 5.7 pp, medium confidence
OSWorld 2.00.0%estimated ± 4.5 pp, low confidence
PostTrainBench v1.119.9%estimated ± 10.5 pp, low confidence
Vibe Code Bench3.5%measured
C-Eval88.9%estimated ± 1.0 pp, low confidence
Claw-Eval57.1%estimated ± 6.8 pp, low confidence
MCP Atlas43.5%estimated ± 14.8 pp, low confidence
MCP-Tasks79.5%estimated ± 8.2 pp, low confidence
MMLU-Redux94.4%estimated ± 1.1 pp, medium confidence
SuperGPQA69.4%estimated ± 7.5 pp, medium confidence
Toolathlon37.5%estimated ± 7.2 pp, medium confidence
HealthBench Hard2.0%estimated ± 7.7 pp, low confidence
MMLU-Pro (Arcee)77.1%estimated ± 4.0 pp, low confidence
React Native Evals30.4%estimated ± 3.0 pp, low confidence
ResearchClawBench14.9%estimated ± 2.2 pp, medium confidence
SWE-Rebench0.0%estimated ± 6.5 pp, low confidence
ARC-AGI-31.7%estimated ± 3.4 pp, medium confidence
BioMysteryBench (human-solvable)81.7%estimated ± 0.6 pp, low confidence
HealthBench (raw)17.6%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.9%estimated ± 3.7 pp, low confidence
HealthBench Professional57.7%estimated ± 4.8 pp, low confidence
HealthBench Professional (raw)30.7%estimated ± 5.0 pp, low confidence
HLE-Verified15.3%estimated ± 2.0 pp, low confidence
LABBench258.2%estimated ± 1.8 pp, low confidence
MMLU88.6%estimated ± 1.1 pp, medium confidence