benchgap
Alibaba (Qwen) · model

Qwen3.7 Plus benchmark scores

As of 2026-10-07, Qwen3.7 Plus (Alibaba (Qwen)) has measured scores on 54 benchmarks and estimated scores on 78 more.

BenchmarkScoreSource
AA-GPQA Diamond90.0%measured
AA-HLE35.6%measured
AA-SciCode46.1%measured
AIME2696.4%estimated ± 1.3 pp, high confidence
Artificial Analysis Intelligence Index25.2%measured
CritPt9.1%measured
GDPval-AA13.5%measured
GPQA Diamond90.3%measured
HLE34.7%measured
IFBench79.1%measured
AA-LCR73.0%measured
LiveCodeBench v687.8%estimated ± 2.1 pp, high confidence
AA-Omniscience Accuracy22.5%measured
SciCode51.3%measured
Terminal-Bench 2.167.8%estimated ± 6.9 pp, medium confidence
BrowseComp82.8%estimated ± 2.2 pp, medium confidence
HLE w/ tools38.5%estimated ± 6.1 pp, low confidence
IFEval94.6%measured
LongBench v261.5%estimated ± 5.9 pp, medium confidence
VITA-Bench45.6%measured
τ²-bench results93.0%measured
AA Coding Index55.9%measured
AA-MMMU-Pro80.5%measured
APEX-Agents28.4%estimated ± 0.6 pp, medium confidence
APEX-Agents-AA22.4%measured
BioMysteryBench (human-difficult)29.3%estimated ± 5.9 pp, low confidence
DeepSearchQA84.1%estimated ± 4.5 pp, high confidence
SWE-bench Verified77.7%measured
AutomationBench28.4%estimated ± 9.7 pp, low confidence
BFCL v472.9%measured
CyberGym60.7%estimated ± 5.3 pp, low confidence
JobBench19.9%estimated ± 4.9 pp, low confidence
SWE-bench Pro57.6%measured
τ³-bench results70.7%estimated ± 1.4 pp, low confidence
WideResearch61.5%estimated ± 9.1 pp, low confidence
AA Agentic Index19.7%measured
MMLU-Pro88.5%measured
FrontierMath v2 (Tiers 1-3)39.0%estimated ± 8.4 pp, low confidence
GPQA90.3%measured
AA-IFBench78.0%measured
Gert Labs47.5%estimated ± 3.7 pp, low confidence
AA-AnalystAgent9.2%estimated ± 5.5 pp, low confidence
AA EnterpriseOps-Gym43.0%estimated ± 3.6 pp, medium confidence
AA Harvey LAB80.8%estimated ± 1.2 pp, low confidence
ApprenticeBench0.0%estimated ± 5.0 pp, low confidence
ARC-AGI-189.6%estimated ± 2.7 pp, high confidence
ARC-AGI-257.0%estimated ± 10.1 pp, low confidence
CursorBench 3.134.4%estimated ± 3.4 pp, low confidence
CursorBench 3.254.7%estimated ± 3.6 pp, high confidence
Terminal-Bench 3.02.7%estimated ± 6.9 pp, low confidence
FrontierCode 1.1 Main17.7%estimated ± 2.7 pp, low confidence
FrontierSWE v210.9%estimated ± 9.0 pp, medium confidence
OfficeQA Pro60.4%estimated ± 7.2 pp, medium confidence
OSWorld-Verified73.3%measured
Vals GPQA Diamond89.6%estimated ± 2.0 pp, high confidence
Vals LiveCodeBench86.6%estimated ± 0.8 pp, medium confidence
Vals MMLU-Pro87.0%estimated ± 1.2 pp, high confidence
Vals SWE-bench75.4%estimated ± 5.3 pp, low confidence
Terminal-Bench 2.1 (Vals AI)52.8%measured
AA AutomationBench3.8%estimated ± 7.0 pp, low confidence
GDP.pdf7.1%estimated ± 4.9 pp, medium confidence
AA ITBench41.3%estimated ± 3.6 pp, low confidence
AA Tau3 Banking22.0%estimated ± 5.8 pp, low confidence
AA Terminal-Bench 2.164.9%estimated ± 4.4 pp, high confidence
AA Terminal-Bench 4.05.6%estimated ± 14.0 pp, low confidence
CursorBench 4.019.7%estimated ± 1.6 pp, medium confidence
CWE-bench v121.3%estimated ± 6.9 pp, low confidence
DeepSWE24.3%estimated ± 6.4 pp, low confidence
HLE w/o tools33.2%estimated ± 2.9 pp, high confidence
OSWorld 2.02.8%measured
PostTrainBench v1.124.2%estimated ± 4.7 pp, medium confidence
SWE Multilingual75.8%measured
Toolathlon-Verified42.5%estimated ± 1.0 pp, low confidence
Vibe Code Bench19.7%estimated ± 8.1 pp, medium confidence
CharXiv85.9%measured
CharXiv w/o tools83.7%estimated ± 1.7 pp, medium confidence
ExploitGym0.6%estimated ± 6.9 pp, low confidence
C-Eval93.1%estimated ± 0.7 pp, medium confidence
Claw-Eval62.7%measured
DeepPlanning62.3%measured
HMMT Feb 202598.4%estimated ± 1.8 pp, low confidence
HMMT Feb 202692.9%measured
MathVision90.3%measured
MCP Atlas73.2%measured
MCP-Tasks77.0%estimated ± 0.6 pp, low confidence
MMAnswerBench86.6%estimated ± 2.8 pp, low confidence
MMLU-ProX85.4%measured
MMLU-Redux94.5%measured
MMMU-Pro79.0%measured
NL2Repo41.1%measured
NOVA-6358.8%measured
QwenClawBench61.8%measured
ScreenSpot Pro79.0%measured
SuperGPQA71.4%measured
Terminal-Bench 2.070.3%measured
Toolathlon27.0%estimated ± 6.3 pp, low confidence
V*97.0%estimated ± 5.1 pp, low confidence
VideoMMMU85.4%measured
ERQA69.8%measured
HealthBench Hard3.5%estimated ± 5.1 pp, low confidence
LiveCodeBench Pro81.8%estimated ± 6.0 pp, low confidence
MedXpertQA (MM)71.0%measured
MedXpertQA (Text)54.7%estimated ± 4.5 pp, medium confidence
MMLU-Pro (Arcee)87.0%estimated ± 3.4 pp, low confidence
React Native Evals78.1%estimated ± 3.5 pp, high confidence
ResearchClawBench17.1%estimated ± 1.3 pp, high confidence
SWE-Rebench61.0%estimated ± 3.6 pp, medium confidence
ARC-AGI-30.0%estimated ± 2.1 pp, medium confidence
BioMysteryBench (human-solvable)82.7%estimated ± 0.6 pp, low confidence
DRACO67.1%estimated ± 3.0 pp, low confidence
FrontierCode 1.1 Extended55.2%estimated ± 3.0 pp, low confidence
HealthBench (raw)44.9%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.9%estimated ± 3.7 pp, low confidence
HealthBench Professional50.7%estimated ± 4.8 pp, low confidence
HealthBench Professional (raw)55.5%estimated ± 5.0 pp, low confidence
HLE-Verified18.0%estimated ± 2.0 pp, low confidence
LABBench277.1%estimated ± 1.8 pp, low confidence
MMMU87.7%estimated ± 0.9 pp, medium confidence
LiveCodeBench89.6%measured
OpenHarmony Bench51.0%estimated ± 2.4 pp, high confidence
IMOAnswerBench86.0%measured
BabyVision63.5%estimated ± 4.3 pp, medium confidence
MMVU77.9%estimated ± 2.9 pp, medium confidence
SimpleVQA81.7%measured
ZeroBench17.2%estimated ± 4.7 pp, low confidence
MMMLU89.0%measured
MMLU87.7%estimated ± 1.1 pp, medium confidence
MMMU-Pro w/ Python80.3%estimated ± 0.6 pp, high confidence
MathVision w/ Python95.3%estimated ± 0.5 pp, medium confidence
RealWorldQA86.9%measured
Video-MME (with subtitle)88.0%measured
OmniDocBench 1.591.4%measured