benchgap
Alibaba (Qwen) · model

Qwen3.8-27B benchmark scores

As of 2026-10-07, Qwen3.8-27B (Alibaba (Qwen)) has measured scores on 44 benchmarks and estimated scores on 79 more.

BenchmarkScoreSource
AA-GPQA Diamond90.5%measured
AA-HLE33.9%measured
AA-SciCode46.6%measured
Artificial Analysis Intelligence Index33.7%measured
CritPt5.4%measured
GDPval-AA46.2%measured
GPQA Diamond89.2%measured
HLE30.8%measured
IFBench79.5%measured
AA-LCR82.0%measured
LiveCodeBench v690.3%measured
AA-Omniscience Accuracy15.6%measured
SciCode44.7%estimated ± 2.9 pp, medium confidence
Terminal-Bench 2.173.0%measured
BrowseComp83.8%estimated ± 3.1 pp, high confidence
HLE w/ tools55.3%estimated ± 6.1 pp, medium confidence
IFEval92.8%estimated ± 3.5 pp, high confidence
LongBench v261.5%estimated ± 5.9 pp, low confidence
τ²-bench results94.3%estimated ± 2.4 pp, medium confidence
AA Coding Index68.1%measured
AA-MMMU-Pro76.3%measured
APEX-Agents38.6%estimated ± 10.0 pp, low confidence
APEX-Agents-AA31.5%estimated ± 3.6 pp, medium confidence
BioMysteryBench (human-difficult)22.7%estimated ± 5.9 pp, low confidence
DeepSearchQA92.9%estimated ± 4.5 pp, high confidence
SWE-bench Verified78.3%estimated ± 2.4 pp, high confidence
AutomationBench38.9%estimated ± 7.0 pp, low confidence
BFCL v477.1%estimated ± 2.6 pp, medium confidence
CyberGym81.3%estimated ± 5.3 pp, medium confidence
JobBench33.4%measured
SWE-bench Pro61.7%measured
WideResearch77.6%estimated ± 1.2 pp, medium confidence
AA Agentic Index46.5%measured
MMLU-Pro85.2%estimated ± 2.5 pp, medium confidence
GPQA89.2%measured
AA-IFBench80.3%estimated ± 8.5 pp, medium confidence
Gert Labs78.5%estimated ± 7.2 pp, low confidence
AA-AnalystAgent34.2%estimated ± 3.6 pp, medium confidence
AA EnterpriseOps-Gym44.2%measured
AA Harvey LAB93.2%estimated ± 1.2 pp, medium confidence
ApprenticeBench29.5%estimated ± 5.2 pp, low confidence
ARC-AGI-187.9%estimated ± 2.7 pp, high confidence
ARC-AGI-252.6%estimated ± 10.1 pp, low confidence
CursorBench 3.135.6%estimated ± 3.4 pp, low confidence
CursorBench 3.264.9%estimated ± 3.3 pp, medium confidence
Terminal-Bench 3.04.9%estimated ± 6.9 pp, low confidence
FrontierCode 1.1 Main24.3%estimated ± 3.7 pp, low confidence
FrontierSWE v218.2%estimated ± 9.0 pp, medium confidence
OfficeQA Pro50.3%estimated ± 7.2 pp, low confidence
OSWorld-Verified84.3%measured
Vals GPQA Diamond88.9%measured
Vals LiveCodeBench84.0%measured
Vals MMLU-Pro84.3%measured
Vals SWE-bench86.0%measured
Terminal-Bench 2.1 (Vals AI)58.4%measured
VulcanBench v382.6%measured
AA AutomationBench48.2%measured
GDP.pdf16.6%measured
AA ITBench50.0%estimated ± 5.7 pp, low confidence
AA Tau3 Banking48.0%measured
AA Terminal-Bench 2.179.8%measured
AA Terminal-Bench 4.05.6%measured
CursorBench 4.030.4%estimated ± 1.6 pp, medium confidence
CWE-bench v146.6%estimated ± 6.0 pp, medium confidence
DeepSWE42.2%measured
HLE w/o tools30.8%measured
OSWorld 2.019.6%estimated ± 3.6 pp, high confidence
PostTrainBench v1.129.1%estimated ± 2.7 pp, high confidence
Terminal-Bench 4.09.5%estimated ± 5.1 pp, low confidence
Terminal-Bench-Science 0.10.0%estimated ± 4.1 pp, low confidence
Toolathlon-Verified73.2%estimated ± 1.3 pp, low confidence
Vibe Code Bench48.0%estimated ± 10.3 pp, low confidence
CharXiv90.2%measured
CharXiv w/o tools83.7%measured
ExploitGym14.3%estimated ± 5.0 pp, low confidence
C-Eval92.9%estimated ± 0.8 pp, medium confidence
Claw-Eval65.0%estimated ± 5.6 pp, low confidence
MathVision90.0%measured
MCP Atlas82.9%estimated ± 2.7 pp, medium confidence
MMLU-Redux91.4%estimated ± 1.1 pp, medium confidence
MMMU-Pro80.3%estimated ± 1.4 pp, medium confidence
NL2Repo42.3%measured
QwenClawBench100.0%estimated ± 6.0 pp, low confidence
ScreenSpot Pro74.5%estimated ± 4.5 pp, medium confidence
SuperGPQA57.4%estimated ± 7.5 pp, medium confidence
Terminal-Bench 2.065.3%estimated ± 5.9 pp, medium confidence
Toolathlon55.9%estimated ± 6.3 pp, medium confidence
V*96.7%estimated ± 5.1 pp, low confidence
VideoMMMU87.2%estimated ± 1.5 pp, high confidence
ERQA65.5%measured
HealthBench Hard0.4%estimated ± 5.1 pp, low confidence
LiveCodeBench Pro85.0%estimated ± 6.0 pp, low confidence
MedXpertQA (MM)76.3%estimated ± 4.2 pp, medium confidence
MedXpertQA (Text)51.5%estimated ± 4.5 pp, medium confidence
MMLU-Pro (Arcee)86.6%estimated ± 3.4 pp, medium confidence
React Native Evals92.0%estimated ± 2.4 pp, low confidence
ResearchClawBench20.7%estimated ± 2.0 pp, medium confidence
SWE-Rebench58.1%estimated ± 8.9 pp, low confidence
ARC-AGI-30.0%estimated ± 2.1 pp, medium confidence
BioMysteryBench (human-solvable)81.8%estimated ± 0.6 pp, low confidence
DRACO80.5%estimated ± 3.0 pp, medium confidence
FrontierCode 1.1 Extended41.2%estimated ± 2.4 pp, low confidence
HealthBench (raw)43.2%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.9%estimated ± 3.7 pp, low confidence
HealthBench Professional51.2%estimated ± 1.3 pp, low confidence
HealthBench Professional (raw)54.0%estimated ± 5.0 pp, low confidence
HLE-Verified21.9%estimated ± 2.0 pp, low confidence
LABBench279.9%estimated ± 1.8 pp, low confidence
MMMU84.1%estimated ± 0.9 pp, high confidence
LiveCodeBench100.0%estimated ± 8.9 pp, low confidence
Agents' Last Exam42.9%measured
OpenHarmony Bench51.0%estimated ± 2.4 pp, high confidence
BabyVision65.7%measured
MMVU77.7%estimated ± 2.9 pp, medium confidence
SimpleVQA67.6%estimated ± 7.4 pp, low confidence
ZeroBench16.9%estimated ± 4.7 pp, low confidence
MMMLU89.5%estimated ± 1.1 pp, medium confidence
MMLU82.6%estimated ± 1.1 pp, medium confidence
MMMU-Pro w/ Python79.5%estimated ± 1.6 pp, high confidence
MathVision w/ Python94.6%measured
RealWorldQA85.9%measured
Video-MME (with subtitle)89.5%estimated ± 0.9 pp, medium confidence
OmniDocBench 1.591.1%measured