benchgap
Zhipu (GLM) · model

GLM-5.3 (max) benchmark scores

As of 2026-10-07, GLM-5.3 (max) (Zhipu (GLM)) has measured scores on 12 benchmarks and estimated scores on 3 more.

BenchmarkScoreSource
Terminal-Bench 2.183.9%measured
Terminal-Bench 4.041.9%measured
Terminal-Bench Science 0.19.5%measured
GPQA Diamond91.7%measured
SciCode59.0%measured
CritPt19.1%measured
GDP.pdf11.2%measured
τ³-Bench Banking50.3%measured
MMMU-Pro75.3%estimated ± 2.6 pp, high confidence
AA-AnalystAgent19.4%estimated ± 3.6 pp, medium confidence
Harvey LAB91.2%estimated ± 1.2 pp, medium confidence
Humanity's Last Exam42.3%measured
AutomationBench62.2%measured
EnterpriseOps-Gym36.4%measured
ITBench SRE46.1%measured