benchgap
Other · model

Granite-4.0-1B benchmark scores

As of 2026-10-10, Granite-4.0-1B (Other) has measured scores on 5 benchmarks and estimated scores on 21 more.

BenchmarkScoreSource
AA-GPQA Diamond29.4%estimated ± 4.7 pp, high confidence
AA-HLE4.9%estimated ± 5.3 pp, medium confidence
Artificial Analysis Intelligence Index4.6%estimated ± 2.0 pp, high confidence
CritPt0.6%estimated ± 3.8 pp, high confidence
GPQA Diamond29.9%estimated ± 0.4 pp, high confidence
HLE0.0%estimated ± 6.4 pp, low confidence
IFBench38.0%estimated ± 9.6 pp, medium confidence
AA-Omniscience Accuracy4.1%estimated ± 4.2 pp, high confidence
IFEval77.4%measured
MMLU-Pro34.0%measured
GPQA29.9%measured
AA-IFBench22.4%estimated ± 7.4 pp, medium confidence
ARC-AGI-10.0%estimated ± 6.9 pp, low confidence
ARC-AGI-20.0%estimated ± 11.6 pp, low confidence
Vals GPQA Diamond13.0%estimated ± 4.8 pp, medium confidence
Vals MMLU-Pro0.0%estimated ± 1.2 pp, medium confidence
HLE w/o tools6.1%estimated ± 4.9 pp, medium confidence
C-Eval3.3%estimated ± 1.3 pp, low confidence
MMLU-Redux57.9%estimated ± 0.6 pp, medium confidence
SuperGPQA36.6%estimated ± 8.2 pp, low confidence
HealthBench Hard0.0%estimated ± 5.1 pp, low confidence
ARC-AGI-30.0%estimated ± 2.1 pp, medium confidence
HealthBench Professional0.1%estimated ± 4.8 pp, low confidence
MMLU59.4%measured
SimpleQA14.4%estimated ± 10.9 pp, low confidence
BBH60.4%measured