benchgap
Zhipu (GLM) · model

GLM-4.6 benchmark scores

As of 2026-10-07, GLM-4.6 (Zhipu (GLM)) has measured scores on 14 benchmarks and estimated scores on 48 more.

BenchmarkScoreSource
AA-GPQA Diamond63.2%measured
AA-HLE5.5%measured
AA-SciCode46.2%estimated ± 3.8 pp, high confidence
AIME2695.5%estimated ± 0.6 pp, medium confidence
Artificial Analysis Intelligence Index14.9%measured
CritPt0.0%measured
GDPval-AA10.7%estimated ± 12.5 pp, low confidence
GPQA Diamond81.4%estimated ± 1.7 pp, medium confidence
HLE19.0%estimated ± 9.5 pp, medium confidence
IFBench55.6%estimated ± 6.9 pp, low confidence
AA-LCR26.3%measured
AA-Omniscience Accuracy21.4%measured
SciCode38.8%estimated ± 2.9 pp, low confidence
IFEval86.9%estimated ± 1.3 pp, high confidence
LongBench v20.0%estimated ± 5.9 pp, low confidence
τ²-bench results76.9%measured
AA Coding Index40.3%estimated ± 6.8 pp, medium confidence
APEX-Agents-AA1.6%estimated ± 11.4 pp, low confidence
BioMysteryBench (human-difficult)28.3%estimated ± 5.9 pp, low confidence
SWE-bench Verified72.2%estimated ± 4.4 pp, high confidence
BFCL v454.5%estimated ± 9.8 pp, low confidence
SWE-bench Pro50.3%estimated ± 2.1 pp, high confidence
AA Agentic Index17.8%estimated ± 12.3 pp, low confidence
MMLU-Pro85.2%estimated ± 3.5 pp, high confidence
FrontierMath v2 (Tier 4)2.1%measured
FrontierMath v2 (Tiers 1-3)3.8%measured
GPQA86.2%estimated ± 1.6 pp, high confidence
AA-IFBench36.7%measured
AA EnterpriseOps-Gym18.8%estimated ± 3.6 pp, low confidence
ARC-AGI-117.6%estimated ± 5.4 pp, low confidence
ARC-AGI-20.0%estimated ± 12.6 pp, low confidence
CursorBench 3.259.1%estimated ± 7.2 pp, low confidence
FrontierCode 1.1 Main1.2%estimated ± 4.3 pp, low confidence
Vals GPQA Diamond74.5%measured
Vals LiveCodeBench81.0%measured
Vals MMLU-Pro82.2%measured
Vals SWE-bench63.8%estimated ± 6.3 pp, medium confidence
VulcanBench v381.4%estimated ± 5.6 pp, low confidence
AA ITBench62.1%estimated ± 3.6 pp, low confidence
HLE w/o tools26.0%estimated ± 5.1 pp, medium confidence
PostTrainBench v1.119.9%estimated ± 10.5 pp, low confidence
Vibe Code Bench3.1%measured
C-Eval88.4%estimated ± 1.0 pp, low confidence
HMMT Nov 202596.7%estimated ± 2.6 pp, medium confidence
MCP-Tasks77.8%estimated ± 8.2 pp, low confidence
MMAnswerBench82.4%estimated ± 1.3 pp, medium confidence
MMLU-Redux93.7%estimated ± 1.1 pp, medium confidence
SuperGPQA66.0%estimated ± 7.5 pp, medium confidence
HealthBench Hard1.6%estimated ± 7.7 pp, low confidence
MMLU-Pro (Arcee)77.1%estimated ± 4.0 pp, low confidence
React Native Evals27.9%estimated ± 3.0 pp, low confidence
SWE-Rebench0.0%estimated ± 6.5 pp, low confidence
ARC-AGI-31.7%estimated ± 3.4 pp, medium confidence
FrontierMath (legacy)4.9%estimated ± 0.7 pp, low confidence
BioMysteryBench (human-solvable)81.7%estimated ± 0.6 pp, low confidence
HealthBench (raw)8.6%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.9%estimated ± 3.7 pp, low confidence
HealthBench Professional49.9%estimated ± 1.3 pp, low confidence
HealthBench Professional (raw)22.6%estimated ± 5.0 pp, low confidence
HLE-Verified15.3%estimated ± 2.0 pp, low confidence
LABBench278.7%estimated ± 1.8 pp, low confidence
MMLU87.1%estimated ± 1.1 pp, medium confidence