benchgap
Other · model

ZAYA1-8B benchmark scores

As of 2026-10-07, ZAYA1-8B (Other) has measured scores on 10 benchmarks and estimated scores on 36 more.

BenchmarkScoreSource
AA-GPQA Diamond69.4%estimated ± 5.4 pp, medium confidence
AA-HLE10.5%estimated ± 4.0 pp, high confidence
AA-SciCode22.1%estimated ± 3.7 pp, high confidence
AIME2689.1%measured
Artificial Analysis Intelligence Index11.2%estimated ± 6.2 pp, medium confidence
CritPt0.6%estimated ± 4.1 pp, high confidence
GDPval-AA0.0%estimated ± 5.8 pp, low confidence
GPQA Diamond71.0%measured
HLE12.7%estimated ± 8.2 pp, low confidence
IFBench52.6%measured
LiveCodeBench v665.8%measured
AA-Omniscience Accuracy16.7%estimated ± 8.1 pp, medium confidence
SciCode27.2%estimated ± 5.4 pp, medium confidence
IFEval85.6%measured
τ²-bench results62.9%estimated ± 14.1 pp, low confidence
AA Coding Index16.5%estimated ± 6.4 pp, medium confidence
SWE-bench Verified59.4%estimated ± 10.5 pp, low confidence
BFCL v439.2%measured
SWE-bench Pro9.3%estimated ± 4.7 pp, medium confidence
MMLU-Pro74.2%measured
FrontierMath v2 (Tiers 1-3)24.4%estimated ± 8.4 pp, low confidence
GPQA71.0%measured
AA-IFBench34.8%estimated ± 8.2 pp, medium confidence
ARC-AGI-149.2%estimated ± 2.7 pp, medium confidence
ARC-AGI-219.8%estimated ± 10.1 pp, low confidence
Vals GPQA Diamond69.7%estimated ± 2.0 pp, medium confidence
Vals MMLU-Pro28.2%estimated ± 1.2 pp, medium confidence
HLE w/o tools6.2%estimated ± 4.4 pp, medium confidence
C-Eval81.5%estimated ± 0.9 pp, low confidence
Claw-Eval45.7%estimated ± 2.3 pp, low confidence
HMMT Feb 202585.5%estimated ± 1.8 pp, low confidence
HMMT Feb 202671.6%measured
MMAnswerBench73.2%estimated ± 1.3 pp, medium confidence
MMLU-Redux86.9%estimated ± 0.7 pp, high confidence
NL2Repo36.2%estimated ± 3.1 pp, low confidence
SuperGPQA47.4%estimated ± 9.2 pp, low confidence
HealthBench Hard0.0%estimated ± 5.1 pp, low confidence
MedXpertQA (Text)23.2%estimated ± 4.5 pp, low confidence
MMLU-Pro (Arcee)79.4%estimated ± 3.4 pp, medium confidence
ARC-AGI-30.0%estimated ± 2.1 pp, medium confidence
HealthBench Professional17.7%estimated ± 4.8 pp, low confidence
AIME 202589.6%estimated ± 1.5 pp, medium confidence
IMOAnswerBench59.3%measured
MMMLU75.7%estimated ± 1.1 pp, low confidence
MMLU89.3%estimated ± 3.6 pp, medium confidence
MATH-50096.8%estimated ± 3.1 pp, medium confidence