benchgap
Alibaba (Qwen) · model

Qwen3.6-35B-A3B benchmark scores

As of 2026-10-07, Qwen3.6-35B-A3B (Alibaba (Qwen)) has measured scores on 47 benchmarks and estimated scores on 83 more.

BenchmarkScoreSource
AA-GPQA Diamond84.1%measured
AA-HLE22.2%measured
AA-SciCode36.6%measured
AIME2692.7%measured
Artificial Analysis Intelligence Index18.2%measured
CritPt0.3%measured
GDPval-AA19.8%measured
GPQA Diamond86.0%estimated ± 0.4 pp, high confidence
HLE21.4%measured
IFBench64.1%estimated ± 6.9 pp, medium confidence
AA-LCR71.7%measured
LiveCodeBench v684.6%estimated ± 2.1 pp, high confidence
AA-Omniscience Accuracy18.8%measured
SciCode36.7%estimated ± 3.2 pp, high confidence
BrowseComp58.9%estimated ± 4.4 pp, medium confidence
HLE w/ tools30.6%estimated ± 2.4 pp, medium confidence
IFEval93.1%estimated ± 1.3 pp, high confidence
LongBench v261.5%estimated ± 5.9 pp, medium confidence
VITA-Bench35.6%measured
τ²-bench results95.3%measured
AA Coding Index41.9%measured
AA-MMMU-Pro75.0%measured
APEX-Agents20.1%estimated ± 0.6 pp, medium confidence
APEX-Agents-AA28.3%estimated ± 5.7 pp, medium confidence
BioMysteryBench (human-difficult)25.9%estimated ± 5.9 pp, low confidence
DeepSearchQA80.0%estimated ± 4.5 pp, high confidence
SWE-bench Verified73.4%measured
AutomationBench28.4%estimated ± 9.7 pp, low confidence
BFCL v472.7%estimated ± 2.6 pp, high confidence
CyberGym57.1%estimated ± 5.3 pp, low confidence
JobBench0.0%estimated ± 5.1 pp, low confidence
SkillsBench18.0%estimated ± 5.6 pp, low confidence
SWE-bench Pro49.5%measured
τ³-bench results67.2%measured
WideResearch60.1%measured
AA Agentic Index15.0%measured
MMLU-Pro85.2%measured
FrontierMath v2 (Tier 4)0.0%estimated ± 3.7 pp, low confidence
FrontierMath v2 (Tiers 1-3)24.4%estimated ± 8.4 pp, low confidence
GPQA86.0%measured
AA-IFBench64.4%measured
Gert Labs42.7%measured
AA-AnalystAgent13.4%estimated ± 12.0 pp, low confidence
AA EnterpriseOps-Gym47.1%estimated ± 3.6 pp, medium confidence
AA Harvey LAB79.7%estimated ± 1.2 pp, low confidence
ApprenticeBench0.0%estimated ± 3.1 pp, low confidence
ARC-AGI-137.9%estimated ± 5.4 pp, low confidence
ARC-AGI-227.2%estimated ± 11.6 pp, low confidence
CursorBench 3.111.0%estimated ± 3.4 pp, low confidence
CursorBench 3.239.1%estimated ± 3.6 pp, medium confidence
FrontierCode 1.1 Main7.7%estimated ± 2.7 pp, low confidence
FrontierSWE v20.0%estimated ± 9.0 pp, low confidence
OfficeQA Pro50.3%estimated ± 7.2 pp, low confidence
OSWorld-Verified55.5%estimated ± 5.9 pp, medium confidence
Vals GPQA Diamond80.3%estimated ± 4.8 pp, high confidence
Vals LiveCodeBench76.5%estimated ± 4.7 pp, medium confidence
Vals MMLU-Pro83.2%estimated ± 1.2 pp, high confidence
Vals SWE-bench59.7%estimated ± 5.4 pp, low confidence
Terminal-Bench 2.1 (Vals AI)45.0%estimated ± 6.3 pp, medium confidence
AA AutomationBench3.8%estimated ± 7.0 pp, medium confidence
GDP.pdf9.4%estimated ± 5.4 pp, medium confidence
AA ITBench32.7%estimated ± 3.6 pp, low confidence
AA Tau3 Banking22.0%estimated ± 5.8 pp, medium confidence
CursorBench 4.011.9%estimated ± 1.6 pp, medium confidence
CWE-bench v129.2%estimated ± 6.9 pp, low confidence
DeepSWE5.4%estimated ± 6.4 pp, low confidence
HLE w/o tools22.9%estimated ± 4.9 pp, high confidence
OSWorld 2.00.0%estimated ± 4.5 pp, low confidence
PostTrainBench v1.119.7%estimated ± 4.7 pp, medium confidence
SWE Multilingual67.2%measured
Toolathlon-Verified40.3%estimated ± 2.3 pp, low confidence
Vibe Code Bench1.4%estimated ± 8.1 pp, medium confidence
CharXiv78.0%measured
CharXiv w/o tools79.3%estimated ± 1.8 pp, medium confidence
ExploitGym0.0%estimated ± 12.2 pp, low confidence
C-Eval90.0%measured
Claw-Eval68.7%measured
DeepPlanning25.9%measured
HMMT Feb 202590.7%measured
HMMT Feb 202683.6%measured
HMMT Nov 202589.1%measured
MathVision85.0%estimated ± 2.6 pp, high confidence
MCP Atlas62.8%measured
MCP-Tasks76.6%estimated ± 0.6 pp, low confidence
MMAnswerBench78.9%measured
MMLU-ProX81.7%estimated ± 1.7 pp, low confidence
MMLU-Redux93.1%estimated ± 0.7 pp, high confidence
MMMU-Pro75.3%measured
NL2Repo29.4%measured
NOVA-6353.6%estimated ± 1.6 pp, low confidence
QwenClawBench52.6%measured
ScreenSpot Pro85.0%estimated ± 2.5 pp, low confidence
SuperGPQA64.7%measured
Terminal-Bench 2.051.5%measured
Toolathlon26.9%measured
V*93.3%estimated ± 1.4 pp, medium confidence
VideoMMMU83.7%measured
ERQA66.7%estimated ± 3.2 pp, medium confidence
HealthBench Hard0.0%estimated ± 5.1 pp, low confidence
LiveCodeBench Pro74.9%estimated ± 6.0 pp, low confidence
MedXpertQA (MM)64.0%estimated ± 2.4 pp, high confidence
MMLU-Pro (Arcee)78.3%estimated ± 4.0 pp, low confidence
React Native Evals74.6%estimated ± 3.5 pp, high confidence
ResearchClawBench18.6%estimated ± 1.3 pp, high confidence
SWE-Rebench57.0%estimated ± 3.6 pp, medium confidence
ARC-AGI-30.0%estimated ± 2.1 pp, medium confidence
BioMysteryBench (human-solvable)81.7%estimated ± 0.6 pp, low confidence
DRACO69.7%estimated ± 3.0 pp, low confidence
FrontierCode 1.1 Extended52.1%estimated ± 3.0 pp, low confidence
HealthBench (raw)30.6%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.9%estimated ± 3.7 pp, low confidence
HealthBench Professional42.4%estimated ± 4.8 pp, low confidence
HealthBench Professional (raw)42.4%estimated ± 5.0 pp, low confidence
HLE-Verified15.3%estimated ± 2.0 pp, low confidence
LABBench268.3%estimated ± 1.8 pp, low confidence
AIME 202593.5%estimated ± 1.5 pp, medium confidence
MMMU81.7%measured
LiveCodeBench80.4%measured
OpenHarmony Bench51.0%estimated ± 2.4 pp, medium confidence
IMOAnswerBench80.1%estimated ± 5.9 pp, medium confidence
SimpleVQA58.9%measured
ZeroBench44.2%estimated ± 6.4 pp, low confidence
MMLU85.4%estimated ± 1.1 pp, medium confidence
MMMU-Pro w/ Python76.8%estimated ± 0.6 pp, high confidence
RefCOCO (avg)92.0%measured
MATH-50097.3%estimated ± 3.1 pp, medium confidence
MathVision w/ Python76.8%estimated ± 0.9 pp, low confidence
RealWorldQA85.3%measured
Video-MME (with subtitle)86.6%measured
OmniDocBench 1.589.9%measured