benchgap
Alibaba (Qwen) · model

Qwen3.6-27B benchmark scores

As of 2026-10-07, Qwen3.6-27B (Alibaba (Qwen)) has measured scores on 44 benchmarks and estimated scores on 85 more.

BenchmarkScoreSource
AA-GPQA Diamond84.2%measured
AA-HLE23.1%measured
AA-SciCode42.8%measured
AIME2694.1%measured
Artificial Analysis Intelligence Index21.4%measured
CritPt1.1%measured
GDPval-AA24.4%measured
GPQA Diamond87.8%estimated ± 0.4 pp, high confidence
HLE24.0%measured
IFBench66.5%estimated ± 6.9 pp, medium confidence
AA-LCR77.3%measured
LiveCodeBench v686.2%estimated ± 2.1 pp, high confidence
AA-Omniscience Accuracy19.6%measured
SciCode42.7%estimated ± 3.2 pp, high confidence
BrowseComp80.0%estimated ± 5.7 pp, medium confidence
HLE w/ tools44.1%estimated ± 6.1 pp, low confidence
IFEval93.3%estimated ± 1.3 pp, high confidence
LongBench v261.5%estimated ± 5.9 pp, medium confidence
VITA-Bench38.4%estimated ± 8.2 pp, low confidence
τ²-bench results94.2%measured
AA Coding Index53.7%measured
AA-MMMU-Pro74.6%measured
APEX-Agents22.4%estimated ± 10.0 pp, low confidence
APEX-Agents-AA20.5%estimated ± 9.1 pp, medium confidence
BioMysteryBench (human-difficult)26.7%estimated ± 5.9 pp, low confidence
DeepSearchQA84.3%estimated ± 4.5 pp, high confidence
SWE-bench Verified77.2%measured
AutomationBench28.4%estimated ± 9.7 pp, low confidence
BFCL v473.9%estimated ± 2.6 pp, high confidence
CyberGym61.0%estimated ± 5.3 pp, low confidence
JobBench42.3%estimated ± 8.4 pp, low confidence
SWE-bench Pro53.5%measured
τ³-bench results68.9%estimated ± 8.9 pp, medium confidence
WideResearch76.5%estimated ± 5.9 pp, low confidence
AA Agentic Index20.1%measured
MMLU-Pro86.2%measured
FrontierMath v2 (Tier 4)0.0%estimated ± 3.7 pp, low confidence
FrontierMath v2 (Tiers 1-3)24.4%estimated ± 8.4 pp, low confidence
GPQA87.8%measured
AA-IFBench67.6%measured
Gert Labs54.8%measured
AA-AnalystAgent17.0%estimated ± 12.0 pp, low confidence
AA EnterpriseOps-Gym45.1%estimated ± 3.6 pp, medium confidence
AA Harvey LAB81.0%estimated ± 1.2 pp, low confidence
ApprenticeBench0.0%estimated ± 3.1 pp, low confidence
ARC-AGI-158.2%estimated ± 5.4 pp, medium confidence
ARC-AGI-238.8%estimated ± 11.6 pp, low confidence
CursorBench 3.126.3%estimated ± 3.4 pp, low confidence
CursorBench 3.247.6%estimated ± 3.6 pp, medium confidence
FrontierCode 1.1 Main16.3%estimated ± 2.7 pp, low confidence
FrontierSWE v23.6%estimated ± 9.0 pp, low confidence
OfficeQA Pro50.3%estimated ± 7.2 pp, low confidence
OSWorld-Verified64.8%estimated ± 5.9 pp, medium confidence
Vals GPQA Diamond83.5%estimated ± 4.8 pp, high confidence
Vals LiveCodeBench80.8%estimated ± 4.7 pp, medium confidence
Vals MMLU-Pro85.2%estimated ± 1.2 pp, high confidence
Vals SWE-bench70.0%measured
Terminal-Bench 2.1 (Vals AI)52.0%estimated ± 6.3 pp, medium confidence
AA AutomationBench4.1%estimated ± 7.0 pp, medium confidence
GDP.pdf9.4%estimated ± 5.4 pp, medium confidence
AA ITBench46.1%estimated ± 3.6 pp, medium confidence
AA Tau3 Banking22.0%estimated ± 5.8 pp, medium confidence
CursorBench 4.018.3%estimated ± 1.6 pp, medium confidence
CWE-bench v134.3%estimated ± 6.9 pp, low confidence
DeepSWE20.2%estimated ± 6.4 pp, low confidence
HLE w/o tools27.6%estimated ± 4.9 pp, high confidence
OSWorld 2.03.4%estimated ± 4.5 pp, low confidence
PostTrainBench v1.120.5%estimated ± 3.3 pp, medium confidence
SWE Multilingual71.3%measured
Toolathlon-Verified48.3%estimated ± 2.3 pp, low confidence
Vibe Code Bench14.7%estimated ± 8.1 pp, medium confidence
CharXiv78.4%measured
CharXiv w/o tools79.9%estimated ± 1.8 pp, medium confidence
ExploitGym0.0%estimated ± 12.2 pp, low confidence
C-Eval91.4%measured
Claw-Eval72.4%measured
DeepPlanning25.3%estimated ± 11.8 pp, low confidence
HMMT Feb 202593.8%measured
HMMT Feb 202684.3%measured
HMMT Nov 202590.7%measured
MathVision88.3%estimated ± 1.4 pp, low confidence
MCP Atlas64.7%estimated ± 7.8 pp, low confidence
MCP-Tasks98.1%estimated ± 6.4 pp, low confidence
MMAnswerBench80.8%measured
MMLU-ProX83.3%estimated ± 1.7 pp, medium confidence
MMLU-Redux93.5%measured
MMMU-Pro75.8%measured
NL2Repo36.2%measured
NOVA-6355.5%estimated ± 1.6 pp, low confidence
QwenClawBench53.4%measured
ScreenSpot Pro85.0%estimated ± 2.5 pp, low confidence
SuperGPQA66.0%measured
Terminal-Bench 2.059.3%measured
Toolathlon40.4%estimated ± 6.3 pp, medium confidence
V*94.7%measured
VideoMMMU84.4%measured
ERQA62.5%measured
HealthBench Hard0.0%estimated ± 5.1 pp, low confidence
LiveCodeBench Pro78.4%estimated ± 6.0 pp, low confidence
MedXpertQA (MM)65.1%estimated ± 2.4 pp, high confidence
MMLU-Pro (Arcee)81.3%estimated ± 4.0 pp, medium confidence
React Native Evals65.7%estimated ± 2.4 pp, low confidence
ResearchClawBench19.5%estimated ± 1.3 pp, high confidence
SWE-Rebench60.2%estimated ± 3.6 pp, medium confidence
ARC-AGI-30.0%estimated ± 2.1 pp, medium confidence
BioMysteryBench (human-solvable)81.7%estimated ± 0.6 pp, low confidence
DRACO71.6%estimated ± 3.0 pp, medium confidence
FrontierCode 1.1 Extended15.8%estimated ± 2.4 pp, low confidence
HealthBench (raw)31.6%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.9%estimated ± 3.7 pp, low confidence
HealthBench Professional45.9%estimated ± 4.8 pp, low confidence
HealthBench Professional (raw)43.3%estimated ± 5.0 pp, low confidence
HLE-Verified15.3%estimated ± 2.0 pp, low confidence
LABBench268.5%estimated ± 1.8 pp, low confidence
AIME 202594.8%estimated ± 1.5 pp, medium confidence
MMMU82.9%measured
LiveCodeBench83.9%measured
OpenHarmony Bench51.0%estimated ± 2.4 pp, medium confidence
IMOAnswerBench80.5%estimated ± 5.9 pp, medium confidence
SimpleVQA56.1%measured
ZeroBench47.9%estimated ± 6.4 pp, low confidence
MMLU85.9%estimated ± 1.1 pp, medium confidence
MMMU-Pro w/ Python77.2%estimated ± 0.6 pp, high confidence
RefCOCO (avg)92.5%measured
MATH-50097.5%estimated ± 3.1 pp, medium confidence
MathVision w/ Python76.8%estimated ± 0.9 pp, low confidence
RealWorldQA84.1%measured
Video-MME (with subtitle)87.7%measured
OmniDocBench 1.591.0%estimated ± 5.9 pp, low confidence