Zhipu (GLM) · model
GLM-4.5 benchmark scores
As of 2026-10-07, GLM-4.5 (Zhipu (GLM)) has measured scores on 3 benchmarks and estimated scores on 23 more.
| Benchmark | Score | Source |
|---|---|---|
| AA-GPQA Diamond | 72.5% | estimated ± 4.2 pp, high confidence |
| AA-HLE | 15.6% | estimated ± 7.4 pp, medium confidence |
| AA-SciCode | 49.6% | estimated ± 4.6 pp, low confidence |
| Artificial Analysis Intelligence Index | 17.2% | estimated ± 7.9 pp, medium confidence |
| CritPt | 2.6% | estimated ± 6.6 pp, medium confidence |
| GPQA Diamond | 79.9% | estimated ± 1.7 pp, medium confidence |
| HLE | 34.1% | estimated ± 9.8 pp, low confidence |
| AA-Omniscience Accuracy | 23.2% | estimated ± 7.7 pp, medium confidence |
| SciCode | 16.4% | estimated ± 2.9 pp, low confidence |
| AA Coding Index | 59.6% | estimated ± 9.0 pp, low confidence |
| SWE-bench Verified | 76.3% | estimated ± 4.4 pp, high confidence |
| SWE-bench Pro | 57.0% | estimated ± 4.8 pp, medium confidence |
| GPQA | 85.2% | estimated ± 1.6 pp, high confidence |
| ARC-AGI-1 | 50.8% | estimated ± 5.9 pp, low confidence |
| ARC-AGI-2 | 0.0% | estimated ± 12.6 pp, low confidence |
| CursorBench 3.2 | 59.1% | estimated ± 7.2 pp, low confidence |
| FrontierCode 1.1 Main | 13.3% | estimated ± 6.2 pp, low confidence |
| Vals GPQA Diamond | 72.2% | measured |
| Vals LiveCodeBench | 67.4% | measured |
| Vals MMLU-Pro | 81.2% | measured |
| VulcanBench v3 | 77.5% | estimated ± 5.6 pp, low confidence |
| HLE w/o tools | 25.6% | estimated ± 5.1 pp, medium confidence |
| ARC-AGI-3 | 0.0% | estimated ± 9.2 pp, low confidence |
| HealthBench Professional | 49.9% | estimated ± 1.3 pp, low confidence |
| HLE-Verified | 0.0% | estimated ± 4.6 pp, low confidence |
| LABBench2 | 78.1% | estimated ± 1.8 pp, low confidence |