benchgap
Alibaba (Qwen) · model

Qwen3.8-Omni-Flash benchmark scores

As of 2026-10-07, Qwen3.8-Omni-Flash (Alibaba (Qwen)) has measured scores on 15 benchmarks and estimated scores on 53 more.

BenchmarkScoreSource
AA-GPQA Diamond88.7%estimated ± 3.4 pp, high confidence
AA-HLE37.9%estimated ± 4.0 pp, high confidence
AA-SciCode50.5%estimated ± 3.0 pp, high confidence
Artificial Analysis Intelligence Index34.9%estimated ± 6.2 pp, medium confidence
CritPt11.6%estimated ± 4.1 pp, high confidence
GPQA Diamond91.0%measured
HLE36.5%measured
IFBench81.5%measured
LiveCodeBench v692.6%measured
AA-Omniscience Accuracy38.3%estimated ± 8.1 pp, medium confidence
SciCode53.0%estimated ± 4.8 pp, high confidence
IFEval92.8%estimated ± 3.5 pp, medium confidence
AA Coding Index72.4%estimated ± 2.5 pp, high confidence
AA-MMMU-Pro79.5%estimated ± 1.9 pp, medium confidence
SWE-bench Verified80.4%estimated ± 2.4 pp, high confidence
SWE-bench Pro63.3%measured
MMLU-Pro86.6%estimated ± 3.0 pp, high confidence
GPQA91.0%measured
AA-IFBench80.3%estimated ± 8.5 pp, medium confidence
ARC-AGI-190.6%estimated ± 2.7 pp, high confidence
ARC-AGI-260.3%estimated ± 10.1 pp, low confidence
CursorBench 3.156.8%estimated ± 5.6 pp, low confidence
CursorBench 3.262.0%estimated ± 3.6 pp, high confidence
FrontierCode 1.1 Main36.9%estimated ± 3.7 pp, low confidence
FrontierSWE v221.0%estimated ± 9.0 pp, medium confidence
Vals GPQA Diamond90.3%estimated ± 2.0 pp, high confidence
Vals LiveCodeBench86.0%estimated ± 4.7 pp, medium confidence
Vals MMLU-Pro87.4%estimated ± 1.2 pp, high confidence
Vals SWE-bench82.9%estimated ± 5.4 pp, medium confidence
CursorBench 4.038.0%estimated ± 3.6 pp, high confidence
DeepSWE57.8%measured
HLE w/o tools35.2%estimated ± 2.9 pp, high confidence
PostTrainBench v1.127.7%estimated ± 4.7 pp, high confidence
SWE Multilingual80.5%measured
CharXiv91.4%measured
CharXiv w/o tools83.5%measured
C-Eval93.2%estimated ± 0.8 pp, low confidence
MathVision91.8%measured
MMLU-ProX86.8%estimated ± 1.7 pp, low confidence
MMLU-Redux95.6%estimated ± 2.1 pp, medium confidence
MMMU-Pro79.9%estimated ± 1.4 pp, medium confidence
NL2Repo48.9%measured
NOVA-6359.4%estimated ± 1.6 pp, low confidence
ScreenSpot Pro77.8%estimated ± 4.5 pp, medium confidence
SuperGPQA76.0%estimated ± 9.1 pp, medium confidence
V*98.8%estimated ± 5.1 pp, low confidence
VideoMMMU87.6%estimated ± 1.5 pp, high confidence
ERQA71.0%measured
HealthBench Hard11.0%estimated ± 5.1 pp, low confidence
LiveCodeBench Pro86.2%estimated ± 6.0 pp, low confidence
MedXpertQA (MM)78.2%estimated ± 4.2 pp, medium confidence
MedXpertQA (Text)56.8%estimated ± 4.5 pp, medium confidence
MMLU-Pro (Arcee)87.2%estimated ± 3.4 pp, low confidence
ARC-AGI-30.0%estimated ± 2.1 pp, medium confidence
FrontierCode 1.1 Extended57.1%estimated ± 3.0 pp, medium confidence
HealthBench Professional52.1%estimated ± 4.8 pp, low confidence
MMMU91.7%estimated ± 6.6 pp, low confidence
OpenHarmony Bench54.7%estimated ± 2.7 pp, high confidence
BabyVision69.7%estimated ± 4.3 pp, medium confidence
MMVU78.7%estimated ± 2.9 pp, medium confidence
SimpleVQA71.2%estimated ± 7.4 pp, low confidence
ZeroBench18.7%estimated ± 4.7 pp, low confidence
MMMLU90.2%estimated ± 1.1 pp, medium confidence
MMLU90.4%estimated ± 3.6 pp, low confidence
MathVision w/ Python96.2%measured
RealWorldQA87.7%measured
Video-MME (with subtitle)89.7%estimated ± 0.9 pp, medium confidence
OmniDocBench 1.591.6%estimated ± 7.4 pp, low confidence