‹ PublicAI Index
The LLM benchmark aggregator.
GLM-5.2
Z.ai · 753B · open weights
Strongest in IT operations (#11 of 11), weakest in Safety (#263 of 337). Above par in 26 of 35 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5.1 and Claude Fable 5 and ahead of Gemini 2.5 Pro and O4 Mini.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Human preference64.3−3.2#26/3421/1
Human preference64.3−3.2#26/3421/1
Coding54.8−14.3#33/1652/5
Code generation54.7−15.3#30/771/2
Agentic coding54.8−13.9#41/1572/4
Repository Q&A54.8−11.9—/0✱✱
Agents58.8−9.3#41/2682/5
Knowledge work61.4−12.2#34/1782/2
Tool use60−14—/81✱0/1
Workflow automation53.2−8.8—/0✱✱
Knowledge55.7−7.8#41/1381/2
Academic knowledge56.9−7#38/1231/1
Factuality53.8−5.2—/0✱✱
Professional52.4−9.2#55/1681/1
IT operations39.6−34.2#11/111/1
Finance55.7−7.2#39/1511/1
Legal54.1−13#44/1511/1
Medical51.8−15#69/1401/1
Reasoning50.2−16.7#90/1784/4
Science56.6−6.9#47/1221/1
Reasoning47.3−21.8#72/1403/3
Mathematics50.9−14.8#78/1401/2
Expert reasoning59.2−1.7—/0✱✱
Core abilities47.3−20.1#123/2042/3
Data analysis50.5−15.6#34/571/1
Language44.2−27.3#38/571/1
Instruction following40.2−33.5#46/571/1
General intelligence50.9−19.2#87/2042/3
Long context60.2−4.4—/0✱✱
Safety46.6−14.9#263/3371/3
Secure code59.7−6.8#68/2741/1
Toxicity avoidance55.1−3.9#105/2721/1
Fairness44.6−25.4#204/3001/2
Harm refusal44.9−16.7#249/3001/2
Jailbreak resistance30.7−34.2#254/2721/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GLM-5.2, left for the other.
§ 3 · Sources
Where the numbers come from
11 publications, 57 figures. Every one links to the page it was read from.
LMArena Text 1476
GDPval-AA 1358
AA-Briefcase 1230
ARC-AGI-2 22.8%
LiveBench 73.2LiveBench · Reasoning 78.6LiveBench · Coding 79.7LiveBench · Agentic Coding 51.8LiveBench · Mathematics 89.8LiveBench · Data Analysis 73.7LiveBench · Language 76.2LiveBench · Instruction Following 62.3
Kagi LLM Benchmark 60%
SimpleBench 58.8%
Vals · Legal Research Bench 31.25%Vals · LegalBench 84.07%Vals · Harvey Legal Agent Benchmark 7.08%Vals · Finance Agent 49.7%Vals · CorpFin 66.12%Vals · TaxEval 73.34%Vals · MedCode 40.77%Vals · MedScribe 83.53%Vals · SRE Bench 0%Vals · SWE-bench Verified 82.8% (Mini-SWE-agent)Vals · Vibe Code Bench 63.96% (OpenHands)Vals · Code Migration 37.87%Vals · GPQA Diamond 85.61%Vals · MMLU Pro 86.71%
Enkrypt · Jailbreak risk 29.1%Enkrypt · Harmful content risk 9.4%Enkrypt · CBRN risk 50.7%Enkrypt · Toxicity risk 2.7%Enkrypt · Bias risk 86.3%Enkrypt · Insecure code risk 13.8%
GDPVal-AA 1498tau3-Banking 34.6%Toolathlon Verified 59.9%Automation Bench Public 26.2%Apex-Agents (pass@1) 26.9%MCPMark 72.4%WildClawBench 55%Terminal-Bench 2.1 77.9%SciCode 50.5%SWE-Atlas-QnA 46.4%SWE Bench Pro 46.7%Humanity's Last Exam (without tools) 41.1%GPQA Diamond 89.5%CritPt 20.9%AA-LCR 76.7%AA-Omniscience Accuracy 24%AA-Omniscience Non-Hallucination 74%
GDPVal-AA (Elo) 1498Toolathlon Verified 59.9%Terminal-Bench 2.1 77.9%SWE-bench Pro (strict) 46.7%MCPMark 72.4%SWE-Atlas-QnA (strict) 46.4%
Badge
[](https://publicai.io/model-index/m/glm-5-2)