benchgap
Alibaba (Qwen) · model

Qwen3.5 Plus benchmark scores

As of 2026-10-07, Qwen3.5 Plus (Alibaba (Qwen)) has measured scores on 4 benchmarks and estimated scores on 28 more.

BenchmarkScoreSource
AA-SciCode48.0%estimated ± 3.8 pp, high confidence
AIME2695.5%estimated ± 0.6 pp, low confidence
GDPval-AA26.7%estimated ± 6.2 pp, medium confidence
BrowseComp66.2%estimated ± 6.2 pp, medium confidence
AA Coding Index47.2%estimated ± 6.8 pp, medium confidence
APEX-Agents-AA22.0%estimated ± 3.6 pp, medium confidence
DeepSearchQA75.7%estimated ± 10.9 pp, low confidence
SWE-bench Verified74.6%estimated ± 4.4 pp, high confidence
CyberGym42.2%estimated ± 11.4 pp, low confidence
JobBench18.5%measured
WideResearch74.4%estimated ± 1.2 pp, medium confidence
FrontierMath v2 (Tier 4)2.1%measured
FrontierMath v2 (Tiers 1-3)21.0%measured
Gert Labs50.4%estimated ± 7.3 pp, medium confidence
ApprenticeBench0.2%estimated ± 5.6 pp, low confidence
FrontierCode 1.1 Main6.3%estimated ± 4.3 pp, low confidence
OSWorld-Verified50.5%estimated ± 8.7 pp, low confidence
Vals SWE-bench70.5%estimated ± 6.3 pp, medium confidence
AA AutomationBench44.7%estimated ± 5.0 pp, low confidence
GDP.pdf17.8%estimated ± 5.2 pp, low confidence
PostTrainBench v1.119.9%estimated ± 10.5 pp, low confidence
Toolathlon-Verified73.2%estimated ± 1.3 pp, low confidence
Vibe Code Bench15.7%measured
Claw-Eval50.7%estimated ± 5.6 pp, low confidence
HMMT Nov 202596.8%estimated ± 2.6 pp, low confidence
MCP Atlas40.3%estimated ± 8.2 pp, medium confidence
MMAnswerBench82.4%estimated ± 1.3 pp, low confidence
Toolathlon30.9%estimated ± 7.4 pp, low confidence
React Native Evals63.9%estimated ± 3.0 pp, low confidence
SWE-Rebench12.4%estimated ± 6.5 pp, low confidence
FrontierMath (legacy)21.9%estimated ± 0.7 pp, low confidence
Agents' Last Exam48.9%estimated ± 8.7 pp, low confidence