benchgap
Alibaba (Qwen) · model

Qwen3.5 Flash benchmark scores

As of 2026-10-07, Qwen3.5 Flash (Alibaba (Qwen)) has measured scores on 6 benchmarks and estimated scores on 36 more.

BenchmarkScoreSource
AA-GPQA Diamond84.1%estimated ± 4.2 pp, high confidence
AA-HLE25.7%estimated ± 7.4 pp, medium confidence
AA-SciCode45.2%estimated ± 4.3 pp, medium confidence
AIME2695.5%estimated ± 0.6 pp, low confidence
Artificial Analysis Intelligence Index22.3%estimated ± 7.9 pp, medium confidence
CritPt4.0%estimated ± 6.6 pp, medium confidence
GPQA Diamond86.6%estimated ± 1.7 pp, medium confidence
HLE39.5%estimated ± 9.8 pp, low confidence
AA-Omniscience Accuracy27.0%estimated ± 7.7 pp, medium confidence
SciCode43.3%estimated ± 2.9 pp, low confidence
AA Coding Index45.3%estimated ± 5.8 pp, medium confidence
SWE-bench Verified76.3%estimated ± 4.4 pp, high confidence
SWE-bench Pro57.0%estimated ± 4.8 pp, high confidence
FrontierMath v2 (Tier 4)0.0%measured
FrontierMath v2 (Tiers 1-3)6.2%measured
GPQA88.0%estimated ± 1.6 pp, high confidence
ARC-AGI-175.1%estimated ± 5.9 pp, medium confidence
ARC-AGI-229.8%estimated ± 12.6 pp, low confidence
CursorBench 3.138.1%estimated ± 5.8 pp, low confidence
CursorBench 3.249.6%estimated ± 5.4 pp, low confidence
FrontierCode 1.1 Main16.6%estimated ± 5.2 pp, low confidence
FrontierSWE v25.0%estimated ± 9.5 pp, low confidence
Vals GPQA Diamond82.8%measured
Vals LiveCodeBench83.3%measured
Vals MMLU-Pro84.1%measured
Vals SWE-bench64.4%measured
VulcanBench v382.6%estimated ± 5.6 pp, low confidence
CursorBench 4.022.2%estimated ± 4.4 pp, low confidence
DeepSWE56.3%estimated ± 7.2 pp, low confidence
HLE w/o tools28.1%estimated ± 5.1 pp, medium confidence
PostTrainBench v1.115.8%estimated ± 2.7 pp, high confidence
HMMT Nov 202599.9%estimated ± 2.6 pp, low confidence
MMAnswerBench81.9%estimated ± 1.3 pp, low confidence
NL2Repo38.9%estimated ± 4.7 pp, medium confidence
React Native Evals57.9%estimated ± 2.4 pp, low confidence
ARC-AGI-30.0%estimated ± 9.2 pp, low confidence
FrontierMath (legacy)7.3%estimated ± 0.7 pp, low confidence
FrontierCode 1.1 Extended10.1%estimated ± 2.4 pp, low confidence
HealthBench Professional50.9%estimated ± 1.3 pp, low confidence
HLE-Verified0.0%estimated ± 4.6 pp, low confidence
LABBench279.8%estimated ± 1.8 pp, low confidence
OpenHarmony Bench49.4%estimated ± 3.2 pp, medium confidence