Request a pilot
‹ PublicAI Index

The LLM benchmark aggregator.

GLM-5.2

Z.ai · 753B · open weights

Strongest in IT operations (#11 of 11), weakest in Safety (#263 of 337). Above par in 26 of 35 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5.1 and Claude Fable 5 and ahead of Gemini 2.5 Pro and O4 Mini.

§ 1 · Profile

What it is good at

Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.

Human preference64.3#26/342
Human preference64.3#26/342
Coding54.8#33/165
Code generation54.7#30/77
Agentic coding54.8#41/157
Repository Q&A54.8—/0✱
Agents58.8#41/268
Knowledge work61.4#34/178
Tool use60—/81✱
Workflow automation53.2—/0✱
Knowledge55.7#41/138
Academic knowledge56.9#38/123
Factuality53.8—/0✱
Professional52.4#55/168
IT operations39.6#11/11
Finance55.7#39/151
Legal54.1#44/151
Medical51.8#69/140
Reasoning50.2#90/178
Science56.6#47/122
Reasoning47.3#72/140
Mathematics50.9#78/140
Expert reasoning59.2—/0✱
Core abilities47.3#123/204
Data analysis50.5#34/57
Language44.2#38/57
Instruction following40.2#46/57
General intelligence50.9#87/204
Long context60.2—/0✱
Safety46.6#263/337
Secure code59.7#68/274
Toxicity avoidance55.1#105/272
Fairness44.6#204/300
Harm refusal44.9#249/300
Jailbreak resistance30.7#254/272

§ 2 · Head to head

What it beats, and what beats it

The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GLM-5.2, left for the other.

GPT-5.4 Miniwon 21 · lost 1
Gemini 3.5 Flash Litewon 21 · lost 1
Command Awon 21 · lost 2
Qwen3.6 Pluswon 19 · lost 3
Nemotron 3 Ultrawon 18 · lost 3
DeepSeek V3won 21 · lost 4
GPT-5.4 Nanowon 19 · lost 4
Grok 4.3won 23 · lost 5
Gemini 2.0 Flashwon 18 · lost 4
GPT-4.1 Nanowon 20 · lost 5
GPT-4.1 Miniwon 20 · lost 5
GPT-4o Miniwon 19 · lost 5
Llama 4 Scout Instructwon 18 · lost 5
Gemma 4 31Bwon 19 · lost 6
Mimo V2.5 Prowon 17 · lost 6
Grok 4.1 Fastwon 16 · lost 6
Gemini 2.5 Flashwon 18 · lost 7
Claude Haiku 4.5won 18 · lost 7
GPT OSS 120Bwon 17 · lost 7
Kimi K2won 17 · lost 7
Claude Sonnet 4won 17 · lost 8
DeepSeek V3.2won 17 · lost 8
Kimi K2.6won 19 · lost 9
GPT OSS 20Bwon 14 · lost 7
GPT-4owon 14 · lost 7
GLM-4.6won 16 · lost 8
O3 Miniwon 15 · lost 8
Inklingwon 22 · lost 12
GPT-5 Nanowon 14 · lost 8
Kimi K2.5won 15 · lost 9
MiniMax M3won 21 · lost 13
Grok 4won 14 · lost 10
DeepSeek R1won 14 · lost 10
DeepSeek V4 Flashwon 18 · lost 13
Claude 3.7 Sonnetwon 12 · lost 9
MiniMax M2.7won 12 · lost 9
Claude Opus 4won 13 · lost 10
GLM-5.1won 12 · lost 10
GPT-4.1won 13 · lost 11
Qwen3 235B A22B Instructwon 13 · lost 11
Qwen3.8 27Bwon 14 · lost 12
Qwen3.7 Maxwon 15 · lost 13
O4 Miniwon 13 · lost 12
Gemini 2.5 Prowon 13 · lost 12
GPT-5 Miniwon 12 · lost 13
GLM-5.3 Flashwon 13 · lost 15
GPT-5.1won 10 · lost 12
GPT-6 Lunawon 11 · lost 14
Gemini 3.6 Flashwon 12 · lost 16
GPT-5.6 Lunawon 14 · lost 19
GPT-5won 10 · lost 14
Gemini 3.5 Flashwon 11 · lost 17
Claude Sonnet 4.6won 10 · lost 16
Claude Sonnet 4.5won 9 · lost 16
GPT-5.2won 10 · lost 18
GPT-5.6 Terrawon 11 · lost 22
DeepSeek V4 Prowon 10 · lost 21
GPT-5.4won 9 · lost 19
Muse Spark 1.1won 9 · lost 19
O3won 8 · lost 17
Grok 4.5won 9 · lost 20
Gemini 3.1 Prowon 7 · lost 16
GPT-6 Solwon 7 · lost 18
Gemini 3.7 Flashwon 6 · lost 17
GLM-5.3won 6 · lost 19
Claude Opus 4.6won 5 · lost 19
Gemini 3.8 Flashwon 5 · lost 23
Claude Opus 4.5won 4 · lost 19
Grok 4.7won 4 · lost 21
DeepSeek V4.1 Flashwon 4 · lost 21
Claude Sonnet 5won 5 · lost 28
Gemini 3 Prowon 3 · lost 18
Grok 4.6won 4 · lost 24
GPT-5.6 Solwon 4 · lost 28
Kimi K3won 3 · lost 22
Qwen3.8 Maxwon 3 · lost 25
GPT-5.5won 3 · lost 26
GPT-6 Astrawon 2 · lost 19
Muse Spark 1.2won 2 · lost 19
Claude Opus 4.8won 2 · lost 26
Claude Opus 5.5won 1 · lost 25
Claude Opus 4.7won 1 · lost 27
Claude Opus 5won 1 · lost 31
Claude Fable 5won 0 · lost 22
Claude Fable 5.1won 0 · lost 23

§ 3 · Sources

Where the numbers come from

11 publications, 57 figures. Every one links to the page it was read from.

LMArena Text ↗1 measureread 2026-09-26
LMArena Text 1476
GDPval-AA ↗1 measureread 2026-09-26
GDPval-AA 1358
AA-Briefcase ↗1 measureread 2026-09-26
AA-Briefcase 1230
ARC-AGI-2 ↗1 measureread 2026-09-26
ARC-AGI-2 22.8%
LiveBench ↗8 measuresread 2026-09-26
LiveBench 73.2LiveBench · Reasoning 78.6LiveBench · Coding 79.7LiveBench · Agentic Coding 51.8LiveBench · Mathematics 89.8LiveBench · Data Analysis 73.7LiveBench · Language 76.2LiveBench · Instruction Following 62.3
Kagi LLM Benchmark ↗1 measureread 2026-09-26
Kagi LLM Benchmark 60%
SimpleBench ↗1 measureread 2026-09-26
SimpleBench 58.8%
Vals.ai ↗14 measuresread 2026-09-26
Vals · Legal Research Bench 31.25%Vals · LegalBench 84.07%Vals · Harvey Legal Agent Benchmark 7.08%Vals · Finance Agent 49.7%Vals · CorpFin 66.12%Vals · TaxEval 73.34%Vals · MedCode 40.77%Vals · MedScribe 83.53%Vals · SRE Bench 0%Vals · SWE-bench Verified 82.8% (Mini-SWE-agent)Vals · Vibe Code Bench 63.96% (OpenHands)Vals · Code Migration 37.87%Vals · GPQA Diamond 85.61%Vals · MMLU Pro 86.71%
Enkrypt AI Safety Leaderboard ↗6 measuresread 2026-09-26
Enkrypt · Jailbreak risk 29.1%Enkrypt · Harmful content risk 9.4%Enkrypt · CBRN risk 50.7%Enkrypt · Toxicity risk 2.7%Enkrypt · Bias risk 86.3%Enkrypt · Insecure code risk 13.8%
GDPVal-AA 1498tau3-Banking 34.6%Toolathlon Verified 59.9%Automation Bench Public 26.2%Apex-Agents (pass@1) 26.9%MCPMark 72.4%WildClawBench 55%Terminal-Bench 2.1 77.9%SciCode 50.5%SWE-Atlas-QnA 46.4%SWE Bench Pro 46.7%Humanity's Last Exam (without tools) 41.1%GPQA Diamond 89.5%CritPt 20.9%AA-LCR 76.7%AA-Omniscience Accuracy 24%AA-Omniscience Non-Hallucination 74%
GDPVal-AA (Elo) 1498Toolathlon Verified 59.9%Terminal-Bench 2.1 77.9%SWE-bench Pro (strict) 46.7%MCPMark 72.4%SWE-Atlas-QnA (strict) 46.4%

Badge

PublicAI Index badge for GLM-5.2[![PublicAI Index](https://publicai.io/model-index/badge?model=glm-5-2)](https://publicai.io/model-index/m/glm-5-2)

✱ Placed by a one-off publication, not a board that re-ran the model. ✱✱ marks a figure the model’s own publisher printed, which counts for less again. Scores are 0–100 on the PublicAI Index scale, 50 = the average of the models each source lists. Scores belong to their publishers.

Snapshot 2026-09-26 · the full index · JSON API