benchgap
Zhipu (GLM) · model

GLM-4.5 benchmark scores

As of 2026-10-07, GLM-4.5 (Zhipu (GLM)) has measured scores on 3 benchmarks and estimated scores on 23 more.

BenchmarkScoreSource
AA-GPQA Diamond72.5%estimated ± 4.2 pp, high confidence
AA-HLE15.6%estimated ± 7.4 pp, medium confidence
AA-SciCode49.6%estimated ± 4.6 pp, low confidence
Artificial Analysis Intelligence Index17.2%estimated ± 7.9 pp, medium confidence
CritPt2.6%estimated ± 6.6 pp, medium confidence
GPQA Diamond79.9%estimated ± 1.7 pp, medium confidence
HLE34.1%estimated ± 9.8 pp, low confidence
AA-Omniscience Accuracy23.2%estimated ± 7.7 pp, medium confidence
SciCode16.4%estimated ± 2.9 pp, low confidence
AA Coding Index59.6%estimated ± 9.0 pp, low confidence
SWE-bench Verified76.3%estimated ± 4.4 pp, high confidence
SWE-bench Pro57.0%estimated ± 4.8 pp, medium confidence
GPQA85.2%estimated ± 1.6 pp, high confidence
ARC-AGI-150.8%estimated ± 5.9 pp, low confidence
ARC-AGI-20.0%estimated ± 12.6 pp, low confidence
CursorBench 3.259.1%estimated ± 7.2 pp, low confidence
FrontierCode 1.1 Main13.3%estimated ± 6.2 pp, low confidence
Vals GPQA Diamond72.2%measured
Vals LiveCodeBench67.4%measured
Vals MMLU-Pro81.2%measured
VulcanBench v377.5%estimated ± 5.6 pp, low confidence
HLE w/o tools25.6%estimated ± 5.1 pp, medium confidence
ARC-AGI-30.0%estimated ± 9.2 pp, low confidence
HealthBench Professional49.9%estimated ± 1.3 pp, low confidence
HLE-Verified0.0%estimated ± 4.6 pp, low confidence
LABBench278.1%estimated ± 1.8 pp, low confidence