benchgap
Alibaba (Qwen) · model

Qwen 3.6 Max (preview) benchmark scores

As of 2026-10-07, Qwen 3.6 Max (preview) (Alibaba (Qwen)) has measured scores on 17 benchmarks and estimated scores on 57 more.

BenchmarkScoreSource
AA-GPQA Diamond88.8%measured
AA-HLE30.8%measured
AA-SciCode47.9%estimated ± 3.0 pp, high confidence
AIME2695.5%estimated ± 0.6 pp, medium confidence
Artificial Analysis Intelligence Index28.4%measured
CritPt3.7%measured
GDPval-AA30.1%estimated ± 12.5 pp, low confidence
GPQA Diamond87.3%estimated ± 2.2 pp, high confidence
HLE35.5%estimated ± 7.8 pp, medium confidence
IFBench75.2%estimated ± 6.9 pp, medium confidence
AA-LCR80.7%measured
LiveCodeBench v688.4%estimated ± 2.1 pp, high confidence
AA-Omniscience Accuracy37.9%measured
SciCode47.0%measured
IFEval93.7%estimated ± 1.3 pp, high confidence
LongBench v261.5%estimated ± 5.9 pp, low confidence
VITA-Bench41.6%estimated ± 8.2 pp, low confidence
τ²-bench results95.9%measured
AA Coding Index61.3%estimated ± 4.6 pp, high confidence
APEX-Agents-AA25.3%estimated ± 11.4 pp, low confidence
BioMysteryBench (human-difficult)39.8%estimated ± 5.9 pp, low confidence
SWE-bench Verified78.5%estimated ± 2.4 pp, high confidence
BFCL v471.5%estimated ± 9.8 pp, low confidence
SWE-bench Pro57.3%measured
AA Agentic Index24.8%estimated ± 12.3 pp, low confidence
MMLU-Pro85.7%estimated ± 3.5 pp, high confidence
FrontierMath v2 (Tier 4)4.2%measured
FrontierMath v2 (Tiers 1-3)23.1%measured
GPQA89.4%estimated ± 2.6 pp, high confidence
AA-IFBench76.6%measured
AA EnterpriseOps-Gym48.1%estimated ± 3.6 pp, medium confidence
ARC-AGI-183.4%estimated ± 5.4 pp, medium confidence
ARC-AGI-246.2%estimated ± 14.2 pp, low confidence
CursorBench 3.152.8%estimated ± 5.6 pp, low confidence
CursorBench 3.254.3%estimated ± 3.6 pp, high confidence
FrontierCode 1.1 Main27.6%estimated ± 5.2 pp, low confidence
FrontierSWE v210.3%estimated ± 9.0 pp, medium confidence
Vals GPQA Diamond86.6%estimated ± 7.4 pp, medium confidence
Vals LiveCodeBench85.6%estimated ± 0.8 pp, medium confidence
Vals MMLU-Pro86.4%estimated ± 2.2 pp, high confidence
Vals SWE-bench72.8%measured
Terminal-Bench 2.1 (Vals AI)57.4%estimated ± 6.3 pp, medium confidence
AA ITBench45.1%estimated ± 3.6 pp, medium confidence
CursorBench 4.033.4%estimated ± 3.6 pp, medium confidence
DeepSWE56.3%estimated ± 7.2 pp, low confidence
HLE w/o tools27.9%estimated ± 5.7 pp, medium confidence
PostTrainBench v1.121.8%estimated ± 3.3 pp, medium confidence
C-Eval94.0%estimated ± 0.7 pp, low confidence
DeepPlanning43.2%estimated ± 11.8 pp, low confidence
HMMT Nov 202593.3%estimated ± 2.6 pp, medium confidence
MCP Atlas71.1%estimated ± 12.3 pp, low confidence
MCP-Tasks65.8%estimated ± 8.2 pp, low confidence
MMAnswerBench82.9%estimated ± 1.3 pp, medium confidence
MMLU-Redux95.6%estimated ± 0.9 pp, medium confidence
NL2Repo42.9%measured
QwenClawBench59.0%measured
SuperGPQA73.9%measured
Terminal-Bench 2.065.4%measured
HealthBench Hard24.3%estimated ± 7.7 pp, low confidence
LiveCodeBench Pro81.6%estimated ± 6.0 pp, low confidence
MMLU-Pro (Arcee)91.0%estimated ± 4.0 pp, low confidence
React Native Evals69.8%estimated ± 2.4 pp, low confidence
ARC-AGI-31.7%estimated ± 3.4 pp, medium confidence
FrontierMath (legacy)23.9%estimated ± 0.7 pp, low confidence
BioMysteryBench (human-solvable)81.7%estimated ± 0.6 pp, low confidence
FrontierCode 1.1 Extended19.4%estimated ± 2.4 pp, low confidence
HealthBench (raw)40.1%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.9%estimated ± 3.7 pp, low confidence
HealthBench Professional57.7%estimated ± 4.8 pp, low confidence
HealthBench Professional (raw)51.1%estimated ± 5.0 pp, low confidence
HLE-Verified15.6%estimated ± 2.0 pp, low confidence
LABBench275.2%estimated ± 1.8 pp, low confidence
OpenHarmony Bench51.6%estimated ± 2.7 pp, high confidence
MMLU92.9%estimated ± 1.1 pp, low confidence