benchgap
Alibaba (Qwen) · model

Qwen3.8-Flash-Next benchmark scores

As of 2026-10-07, Qwen3.8-Flash-Next (Alibaba (Qwen)) has measured scores on 30 benchmarks and estimated scores on 85 more.

BenchmarkScoreSource
AA-GPQA Diamond92.3%measured
AA-HLE38.0%measured
AA-SciCode50.6%measured
Artificial Analysis Intelligence Index39.8%measured
CritPt11.1%measured
GDPval-AA56.6%measured
GPQA Diamond91.7%measured
HLE35.9%measured
IFBench81.3%measured
AA-LCR79.7%measured
LiveCodeBench v691.9%measured
AA-Omniscience Accuracy24.5%measured
SciCode50.3%estimated ± 3.2 pp, high confidence
BrowseComp82.8%estimated ± 2.2 pp, high confidence
HLE w/ tools58.0%estimated ± 5.9 pp, medium confidence
IFEval92.8%estimated ± 3.5 pp, medium confidence
LongBench v261.5%estimated ± 5.9 pp, low confidence
AA Coding Index73.1%measured
AA-MMMU-Pro79.8%measured
APEX-Agents45.3%estimated ± 10.0 pp, low confidence
APEX-Agents-AA36.3%estimated ± 2.9 pp, medium confidence
BioMysteryBench (human-difficult)30.9%estimated ± 5.9 pp, low confidence
DeepSearchQA90.4%estimated ± 8.4 pp, medium confidence
SWE-bench Verified80.2%estimated ± 2.4 pp, high confidence
AutomationBench44.5%estimated ± 9.7 pp, low confidence
BFCL v477.9%estimated ± 2.6 pp, medium confidence
CyberGym100.0%estimated ± 5.4 pp, low confidence
JobBench55.7%measured
SWE-bench Pro62.5%measured
WideResearch82.5%estimated ± 1.2 pp, medium confidence
AA Agentic Index51.4%estimated ± 2.8 pp, high confidence
MMLU-Pro85.5%estimated ± 2.5 pp, medium confidence
GPQA91.7%measured
AA-IFBench80.3%estimated ± 8.5 pp, medium confidence
Gert Labs71.9%estimated ± 3.7 pp, medium confidence
AA-AnalystAgent53.9%estimated ± 5.5 pp, medium confidence
AA EnterpriseOps-Gym47.0%estimated ± 6.7 pp, low confidence
AA Harvey LAB89.7%estimated ± 1.3 pp, medium confidence
ApprenticeBench19.0%estimated ± 5.6 pp, low confidence
ARC-AGI-191.6%estimated ± 2.7 pp, high confidence
ARC-AGI-263.8%estimated ± 10.1 pp, low confidence
CursorBench 3.145.5%estimated ± 3.4 pp, low confidence
CursorBench 3.261.1%estimated ± 3.6 pp, high confidence
FrontierCode 1.1 Main37.7%estimated ± 3.7 pp, low confidence
FrontierSWE v219.6%estimated ± 9.0 pp, medium confidence
OfficeQA Pro54.7%estimated ± 7.2 pp, medium confidence
OSWorld-Verified82.8%estimated ± 3.0 pp, high confidence
Vals GPQA Diamond91.0%estimated ± 2.0 pp, high confidence
Vals LiveCodeBench85.8%estimated ± 4.7 pp, medium confidence
Vals MMLU-Pro87.7%estimated ± 1.2 pp, high confidence
Vals SWE-bench82.1%estimated ± 5.4 pp, medium confidence
AA AutomationBench55.2%estimated ± 5.0 pp, medium confidence
GDP.pdf24.4%estimated ± 3.4 pp, high confidence
AA Tau3 Banking47.5%estimated ± 5.8 pp, medium confidence
CursorBench 4.036.5%estimated ± 1.6 pp, high confidence
CWE-bench v160.0%estimated ± 6.9 pp, medium confidence
DeepSWE58.7%measured
HLE w/o tools35.9%measured
OSWorld 2.019.4%measured
PostTrainBench v1.127.2%estimated ± 4.7 pp, high confidence
SWE Multilingual81.0%measured
Toolathlon-Verified73.5%measured
Vibe Code Bench61.2%estimated ± 10.3 pp, low confidence
CharXiv90.6%measured
CharXiv w/o tools84.6%measured
ExploitGym5.2%estimated ± 6.9 pp, low confidence
C-Eval93.1%estimated ± 0.8 pp, low confidence
Claw-Eval66.2%estimated ± 5.6 pp, low confidence
MathVision90.6%measured
MCP Atlas80.9%estimated ± 7.8 pp, medium confidence
MMLU-ProX87.0%estimated ± 1.7 pp, low confidence
MMLU-Redux94.5%estimated ± 1.1 pp, medium confidence
MMMU-Pro81.5%estimated ± 1.4 pp, medium confidence
NL2Repo48.1%measured
NOVA-6359.6%estimated ± 1.6 pp, low confidence
QwenClawBench100.0%estimated ± 6.0 pp, low confidence
ScreenSpot Pro75.6%estimated ± 4.5 pp, medium confidence
SuperGPQA69.5%estimated ± 7.5 pp, medium confidence
Toolathlon55.9%estimated ± 6.3 pp, low confidence
V*97.4%estimated ± 5.1 pp, low confidence
VideoMMMU87.3%estimated ± 1.5 pp, high confidence
ERQA72.3%measured
HealthBench Hard23.3%estimated ± 5.1 pp, low confidence
LiveCodeBench Pro85.6%estimated ± 6.0 pp, low confidence
MedXpertQA (MM)78.5%estimated ± 4.2 pp, medium confidence
MedXpertQA (Text)59.2%estimated ± 4.5 pp, medium confidence
MMLU-Pro (Arcee)87.4%estimated ± 3.4 pp, low confidence
React Native Evals82.4%estimated ± 3.5 pp, high confidence
ResearchClawBench19.9%estimated ± 2.2 pp, low confidence
SWE-Rebench58.1%estimated ± 8.9 pp, low confidence
ARC-AGI-30.0%estimated ± 2.1 pp, medium confidence
BioMysteryBench (human-solvable)84.3%estimated ± 0.6 pp, low confidence
DRACO84.8%estimated ± 3.0 pp, medium confidence
FrontierCode 1.1 Extended56.8%estimated ± 3.0 pp, low confidence
HealthBench (raw)47.2%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.9%estimated ± 3.7 pp, medium confidence
HealthBench Professional53.5%estimated ± 4.8 pp, low confidence
HealthBench Professional (raw)57.7%estimated ± 5.0 pp, low confidence
HLE-Verified49.9%estimated ± 2.0 pp, medium confidence
LABBench280.7%estimated ± 1.8 pp, low confidence
MMMU87.1%estimated ± 0.9 pp, medium confidence
LiveCodeBench100.0%estimated ± 8.9 pp, low confidence
Agents' Last Exam51.2%measured
OpenHarmony Bench56.4%estimated ± 2.4 pp, high confidence
BabyVision64.7%estimated ± 4.3 pp, medium confidence
MMVU78.1%estimated ± 2.9 pp, medium confidence
SimpleVQA72.1%estimated ± 7.4 pp, low confidence
ZeroBench17.5%estimated ± 4.7 pp, low confidence
MMMLU90.4%estimated ± 1.1 pp, medium confidence
MMLU88.7%estimated ± 1.1 pp, medium confidence
MMMU-Pro w/ Python81.6%estimated ± 1.6 pp, high confidence
MathVision w/ Python95.7%measured
RealWorldQA88.5%measured
Video-MME (with subtitle)89.5%estimated ± 0.9 pp, medium confidence
OmniDocBench 1.592.0%estimated ± 6.4 pp, low confidence