benchgap
Other · model

LLaDA2.2-flash benchmark scores

As of 2026-10-07, LLaDA2.2-flash (Other) has measured scores on 7 benchmarks and estimated scores on 43 more.

BenchmarkScoreSource
AA-SciCode34.7%estimated ± 3.8 pp, high confidence
GDPval-AA3.0%estimated ± 5.8 pp, medium confidence
LiveCodeBench v676.5%estimated ± 2.1 pp, high confidence
SciCode33.0%estimated ± 3.8 pp, high confidence
BrowseComp59.0%estimated ± 8.4 pp, medium confidence
HLE w/ tools23.9%estimated ± 7.2 pp, low confidence
VITA-Bench32.6%estimated ± 9.9 pp, low confidence
τ²-bench results80.3%measured
AA Coding Index23.8%estimated ± 5.1 pp, medium confidence
APEX-Agents6.8%estimated ± 0.6 pp, low confidence
APEX-Agents-AA20.4%estimated ± 5.7 pp, medium confidence
SWE-bench Verified49.3%measured
BFCL v460.8%measured
CyberGym58.4%estimated ± 9.9 pp, medium confidence
JobBench22.7%estimated ± 6.8 pp, medium confidence
SWE-bench Pro30.1%measured
τ³-bench results69.3%estimated ± 1.8 pp, medium confidence
AA Agentic Index20.6%estimated ± 5.1 pp, medium confidence
Gert Labs53.2%estimated ± 8.2 pp, low confidence
AA EnterpriseOps-Gym23.1%estimated ± 3.6 pp, low confidence
CursorBench 3.131.7%estimated ± 5.6 pp, low confidence
CursorBench 3.23.9%estimated ± 3.6 pp, medium confidence
FrontierCode 1.1 Main0.0%estimated ± 2.7 pp, low confidence
FrontierSWE v20.0%estimated ± 9.0 pp, low confidence
OSWorld-Verified60.2%estimated ± 10.9 pp, low confidence
Vals LiveCodeBench20.8%estimated ± 4.7 pp, low confidence
Vals SWE-bench34.1%estimated ± 6.8 pp, low confidence
AA AutomationBench0.0%estimated ± 13.0 pp, low confidence
AA ITBench21.5%estimated ± 3.6 pp, low confidence
AA Tau3 Banking8.2%estimated ± 11.8 pp, low confidence
CursorBench 4.012.6%estimated ± 3.6 pp, medium confidence
PostTrainBench v1.110.6%estimated ± 4.7 pp, medium confidence
SWE Multilingual25.0%measured
Toolathlon-Verified31.4%estimated ± 9.3 pp, low confidence
Vibe Code Bench0.6%estimated ± 8.1 pp, low confidence
Claw-Eval64.2%measured
DeepPlanning32.3%estimated ± 13.1 pp, low confidence
MCP Atlas46.2%measured
MCP-Tasks73.8%estimated ± 0.6 pp, medium confidence
MMLU-ProX63.2%estimated ± 1.7 pp, low confidence
NL2Repo5.4%estimated ± 6.1 pp, low confidence
NOVA-6326.8%estimated ± 1.6 pp, low confidence
QwenClawBench53.7%estimated ± 2.6 pp, high confidence
Toolathlon37.2%estimated ± 7.5 pp, medium confidence
LiveCodeBench Pro54.0%estimated ± 6.0 pp, low confidence
ResearchClawBench17.5%estimated ± 1.3 pp, high confidence
SWE-Rebench57.0%estimated ± 3.6 pp, low confidence
FrontierCode 1.1 Extended41.5%estimated ± 3.0 pp, low confidence
LiveCodeBench63.9%estimated ± 10.2 pp, low confidence
OpenHarmony Bench53.3%estimated ± 4.9 pp, low confidence