benchgap
Other · model

LLaDA2.2-mini benchmark scores

As of 2026-10-07, LLaDA2.2-mini (Other) has measured scores on 8 benchmarks and estimated scores on 46 more.

BenchmarkScoreSource
AA-GPQA Diamond41.8%estimated ± 5.6 pp, low confidence
AA-HLE0.0%estimated ± 4.0 pp, medium confidence
AA-SciCode0.2%estimated ± 3.7 pp, medium confidence
AIME2635.1%measured
Artificial Analysis Intelligence Index0.6%estimated ± 7.4 pp, low confidence
CritPt0.0%estimated ± 4.9 pp, medium confidence
GDPval-AA0.0%estimated ± 5.8 pp, low confidence
GPQA Diamond44.4%measured
HLE0.0%estimated ± 9.2 pp, low confidence
IFBench24.9%measured
AA-LCR56.5%estimated ± 7.0 pp, low confidence
LiveCodeBench v628.1%measured
AA-Omniscience Accuracy4.1%estimated ± 8.6 pp, low confidence
SciCode8.0%estimated ± 5.4 pp, low confidence
BrowseComp67.7%estimated ± 9.4 pp, low confidence
IFEval78.3%estimated ± 3.5 pp, medium confidence
LongBench v235.0%measured
VITA-Bench35.0%estimated ± 10.7 pp, low confidence
τ²-bench results57.5%measured
AA Coding Index0.1%estimated ± 6.4 pp, low confidence
APEX-Agents-AA1.6%estimated ± 11.4 pp, low confidence
SWE-bench Verified30.6%estimated ± 10.5 pp, low confidence
BFCL v447.7%measured
JobBench24.8%estimated ± 8.4 pp, low confidence
SWE-bench Pro0.0%estimated ± 4.7 pp, medium confidence
AA Agentic Index17.5%estimated ± 5.1 pp, medium confidence
MMLU-Pro55.2%estimated ± 5.4 pp, medium confidence
GPQA44.4%estimated ± 0.4 pp, high confidence
AA-IFBench56.8%estimated ± 8.5 pp, low confidence
Gert Labs50.3%estimated ± 9.1 pp, low confidence
AA EnterpriseOps-Gym4.7%estimated ± 3.6 pp, low confidence
ARC-AGI-14.8%estimated ± 2.7 pp, medium confidence
ARC-AGI-27.0%estimated ± 10.1 pp, low confidence
Vals GPQA Diamond42.4%estimated ± 2.0 pp, medium confidence
Vals MMLU-Pro54.8%estimated ± 1.6 pp, medium confidence
AA ITBench87.3%estimated ± 3.6 pp, low confidence
HLE w/o tools5.7%estimated ± 4.4 pp, medium confidence
Claw-Eval57.2%measured
HMMT Feb 20251.1%estimated ± 1.9 pp, low confidence
HMMT Feb 202624.2%estimated ± 4.0 pp, medium confidence
MCP Atlas45.8%estimated ± 11.2 pp, low confidence
MCP-Tasks68.1%estimated ± 6.4 pp, low confidence
MMAnswerBench18.8%estimated ± 1.3 pp, medium confidence
MMLU-Redux78.1%estimated ± 3.0 pp, medium confidence
NL2Repo17.3%estimated ± 3.1 pp, low confidence
SuperGPQA23.7%estimated ± 13.9 pp, low confidence
MedXpertQA (Text)8.9%estimated ± 4.5 pp, low confidence
MMLU-Pro (Arcee)63.9%estimated ± 3.4 pp, low confidence
ResearchClawBench15.7%estimated ± 1.3 pp, high confidence
ARC-AGI-30.0%estimated ± 3.8 pp, medium confidence
HealthBench Professional1.2%estimated ± 4.8 pp, low confidence
AIME 202539.1%estimated ± 1.5 pp, low confidence
MMMLU17.8%estimated ± 1.1 pp, low confidence
MATH-50089.5%estimated ± 3.1 pp, low confidence