benchgap
Other · model

Ling 3.1 Flash benchmark scores

As of 2026-10-07, Ling 3.1 Flash (Other) has measured scores on 13 benchmarks and estimated scores on 67 more.

BenchmarkScoreSource
AA-GPQA Diamond96.2%estimated ± 2.1 pp, low confidence
AA-HLE39.4%measured
AA-SciCode54.1%measured
Artificial Analysis Intelligence Index41.1%measured
CritPt18.0%measured
GDPval-AA56.1%measured
GPQA Diamond96.0%estimated ± 1.0 pp, low confidence
HLE50.4%estimated ± 9.5 pp, medium confidence
AA-LCR83.0%measured
LiveCodeBench v695.1%estimated ± 3.3 pp, medium confidence
AA-Omniscience Accuracy29.1%measured
SciCode53.7%estimated ± 3.2 pp, high confidence
BrowseComp88.9%estimated ± 5.2 pp, low confidence
HLE w/ tools65.4%estimated ± 5.0 pp, low confidence
LongBench v261.5%estimated ± 5.9 pp, low confidence
AA Coding Index70.4%estimated ± 6.4 pp, medium confidence
APEX-Agents45.0%estimated ± 10.0 pp, low confidence
APEX-Agents-AA39.5%estimated ± 6.0 pp, low confidence
BioMysteryBench (human-difficult)34.3%estimated ± 5.9 pp, low confidence
DeepSearchQA90.4%estimated ± 8.4 pp, medium confidence
SWE-bench Verified87.2%estimated ± 6.6 pp, medium confidence
AutomationBench52.5%measured
BFCL v477.9%estimated ± 2.6 pp, medium confidence
CyberGym87.9%measured
JobBench59.5%estimated ± 8.6 pp, medium confidence
SkillsBench68.7%measured
SWE-bench Pro65.6%estimated ± 6.8 pp, medium confidence
WideResearch82.7%estimated ± 1.5 pp, medium confidence
AA Agentic Index52.0%estimated ± 1.7 pp, high confidence
MMLU-Pro85.7%estimated ± 3.5 pp, high confidence
GPQA96.0%estimated ± 1.0 pp, low confidence
Gert Labs70.8%estimated ± 6.0 pp, low confidence
AA-AnalystAgent47.6%estimated ± 12.0 pp, low confidence
AA EnterpriseOps-Gym31.7%estimated ± 5.1 pp, low confidence
ApprenticeBench30.9%estimated ± 4.0 pp, low confidence
ARC-AGI-193.4%estimated ± 5.4 pp, medium confidence
ARC-AGI-274.9%estimated ± 14.2 pp, low confidence
CursorBench 3.154.1%estimated ± 3.4 pp, medium confidence
CursorBench 3.260.2%estimated ± 5.3 pp, medium confidence
FrontierCode 1.1 Main42.7%estimated ± 6.6 pp, medium confidence
FrontierSWE v224.3%estimated ± 14.6 pp, low confidence
OSWorld-Verified80.7%estimated ± 4.7 pp, high confidence
Vals GPQA Diamond98.2%estimated ± 1.5 pp, low confidence
Vals MMLU-Pro93.8%estimated ± 1.5 pp, low confidence
Vals SWE-bench84.4%estimated ± 7.6 pp, low confidence
Terminal-Bench 2.1 (Vals AI)77.7%estimated ± 4.6 pp, medium confidence
AA AutomationBench64.4%estimated ± 7.0 pp, medium confidence
GDP.pdf22.5%estimated ± 5.4 pp, medium confidence
AA Tau3 Banking47.5%estimated ± 5.8 pp, medium confidence
AA Terminal-Bench 4.037.1%estimated ± 4.7 pp, high confidence
CursorBench 4.037.1%estimated ± 3.5 pp, high confidence
CWE-bench v159.7%estimated ± 6.9 pp, medium confidence
HLE w/o tools37.7%estimated ± 5.7 pp, medium confidence
PostTrainBench v1.130.3%estimated ± 7.2 pp, low confidence
Terminal-Bench 4.040.4%measured
Toolathlon-Verified74.3%estimated ± 2.2 pp, low confidence
ExploitGym17.8%estimated ± 9.3 pp, low confidence
C-Eval95.8%estimated ± 1.0 pp, low confidence
MCP Atlas84.2%estimated ± 11.1 pp, low confidence
MMLU-Redux95.3%estimated ± 1.1 pp, medium confidence
NL2Repo56.3%estimated ± 6.5 pp, medium confidence
QwenClawBench100.0%estimated ± 6.0 pp, low confidence
SuperGPQA73.8%estimated ± 7.5 pp, medium confidence
Toolathlon55.9%estimated ± 6.3 pp, low confidence
HealthBench Hard33.7%estimated ± 7.7 pp, low confidence
MMLU-Pro (Arcee)100.0%estimated ± 4.0 pp, low confidence
React Native Evals78.4%estimated ± 4.7 pp, medium confidence
ResearchClawBench19.9%estimated ± 2.2 pp, low confidence
ARC-AGI-31.7%estimated ± 3.4 pp, medium confidence
BioMysteryBench (human-solvable)88.8%estimated ± 0.6 pp, medium confidence
DRACO85.5%measured
HealthBench (raw)48.5%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)60.6%estimated ± 2.2 pp, medium confidence
HealthBench Professional65.4%measured
HealthBench Professional (raw)58.9%estimated ± 5.0 pp, medium confidence
HLE-Verified23.3%estimated ± 6.0 pp, low confidence
LABBench279.5%estimated ± 2.4 pp, low confidence
Agents' Last Exam30.3%estimated ± 2.1 pp, high confidence
OpenHarmony Bench58.8%estimated ± 2.4 pp, high confidence
MMLU90.5%estimated ± 1.1 pp, medium confidence