Request a pilot
‹ PublicAI Index

The LLM benchmark aggregator.

GLM-5.3

Z.ai · 753B · open weights

Strongest in Agentic coding (#8 of 157, on 3 of its 4 boards), weakest in Mathematics (#86 of 140). Above par in 23 of 26 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5.1 and Muse Spark 1.3 and ahead of Claude Opus 4.5 and Claude Sonnet 5.

§ 1 · Profile

What it is good at

Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.

Coding60.1#9/165
Agentic coding61.9#8/157
Code generation53.3#32/77
Repository Q&A46—/0✱
Agents63.8#11/268
Knowledge work67.9#10/178
Workflow automation53—/0✱
Human preference64.7#20/342
Human preference64.7#20/342
Professional56.6#21/168
Legal59.3#15/151
Medical55.5#35/140
Finance55.7#37/151
Cybersecurity51.9—/9✱
Knowledge55.8#40/138
Academic knowledge56.9#37/123
Core abilities53.4#58/204
Language51.2#27/57
Instruction following52.4#28/57
General intelligence59.3#32/204
Data analysis44.6#44/57
Reasoning53.2#60/178
Reasoning56.1#35/140
Science58.4#36/122
Mathematics48#86/140
Expert reasoning53.1—/0✱

§ 2 · Head to head

What it beats, and what beats it

The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GLM-5.3, left for the other.

Llama 4 Scout Instructwon 16 · lost 0
Command Awon 16 · lost 0
Gemini 2.0 Flashwon 16 · lost 0
Inkling Smallwon 16 · lost 0
Mistral Medium 3.5won 17 · lost 0
Mimo V2.5 Prowon 17 · lost 0
Nemotron 3 Ultrawon 18 · lost 0
GPT-4o Miniwon 18 · lost 0
GPT-4.1 Nanowon 18 · lost 0
GPT-4.1 Miniwon 18 · lost 0
Gemini 2.5 Flashwon 18 · lost 0
Mistral Large 3won 18 · lost 0
DeepSeek V3won 19 · lost 0
Gemini 3.5 Flash Litewon 22 · lost 0
GPT-5.4 Nanowon 21 · lost 1
Grok 4.3won 21 · lost 1
GPT-5.4 Miniwon 20 · lost 1
GPT OSS 120Bwon 17 · lost 1
Claude Haiku 4.5won 17 · lost 1
Kimi K2won 16 · lost 1
GLM-4.6won 16 · lost 1
Qwen3.6 27Bwon 15 · lost 1
MiniMax M2.5won 15 · lost 1
GLM-4.7won 15 · lost 1
GPT-6 Lunawon 17 · lost 2
GPT-5 Miniwon 16 · lost 2
DeepSeek V3.2won 16 · lost 2
GPT-4.1won 15 · lost 2
Claude Opus 4won 15 · lost 2
Inklingwon 22 · lost 3
Grok 3 Miniwon 14 · lost 2
Grok 4 Fastwon 14 · lost 2
Gemma 4 31Bwon 14 · lost 2
Qwen3.6 Pluswon 19 · lost 3
Kimi K2.6won 19 · lost 3
Claude Sonnet 4won 16 · lost 3
Gemini 2.5 Prowon 16 · lost 3
O4 Miniwon 15 · lost 3
DeepSeek R1won 15 · lost 3
O3won 15 · lost 3
Kimi K2.5won 15 · lost 3
Qwen3.8 27Bwon 19 · lost 4
O3 Miniwon 14 · lost 3
Qwen3 235B A22B Instructwon 14 · lost 3
Gemini 3.1 Flash Litewon 14 · lost 3
GLM-5.3 Flashwon 18 · lost 4
DeepSeek V4 Flashwon 21 · lost 5
MiniMax M3won 20 · lost 5
Claude Sonnet 4.5won 14 · lost 4
Qwen3.7 Maxwon 17 · lost 5
GLM-5.2won 19 · lost 6
GLM-5won 12 · lost 4
GLM-5.1won 12 · lost 4
GPT-5.6 Lunawon 17 · lost 7
Grok 4won 12 · lost 5
GPT-5.1won 11 · lost 5
Claude Sonnet 4.6won 15 · lost 7
GPT-5won 12 · lost 6
GPT-5.2won 14 · lost 7
DeepSeek V4 Prowon 17 · lost 9
Gemini 3.5 Flashwon 13 · lost 9
GPT-5.6 Terrawon 14 · lost 10
GPT-6 Solwon 11 · lost 8
Gemini 3.6 Flashwon 12 · lost 10
Claude Sonnet 5won 13 · lost 11
Claude Opus 4.5won 11 · lost 10
Grok 4.5won 11 · lost 11
GPT-5.4won 11 · lost 11
Muse Spark 1.1won 10 · lost 12
Gemini 3.8 Flashwon 10 · lost 13
Grok 4.7won 8 · lost 11
Qwen3.8 Maxwon 9 · lost 13
Gemini 3.1 Prowon 9 · lost 13
Gemini 3.7 Flashwon 9 · lost 13
GPT-6 Astrawon 8 · lost 12
DeepSeek V4.1 Flashwon 10 · lost 15
Claude Opus 4.7won 8 · lost 14
Claude Opus 4.6won 7 · lost 13
GPT-5.6 Solwon 9 · lost 17
Muse Spark 1.2won 7 · lost 14
Claude Opus 4.8won 7 · lost 15
Gemini 3 Prowon 5 · lost 12
GPT-5.5won 6 · lost 16
Kimi K3won 7 · lost 19
Grok 4.6won 6 · lost 17
Claude Opus 5won 3 · lost 23
Claude Opus 5.5won 2 · lost 17
Claude Fable 5won 1 · lost 21
Muse Spark 1.3won 0 · lost 19
Claude Fable 5.1won 0 · lost 23

§ 3 · Sources

Where the numbers come from

9 publications, 39 figures. Every one links to the page it was read from.

LMArena Text ↗1 measureread 2026-09-26
LMArena Text 1480
Artificial Analysis ↗1 measureread 2026-09-26
Artificial Analysis Intelligence Index 45
GDPval-AA ↗1 measureread 2026-09-26
GDPval-AA 1646
AA-Briefcase ↗1 measureread 2026-09-26
AA-Briefcase 1516
Terminal-Bench ↗1 measureread 2026-09-26
Terminal-Bench 41.8% (Claude Code)
LiveBench ↗8 measuresread 2026-09-26
LiveBench 76.1LiveBench · Reasoning 85.8LiveBench · Coding 79LiveBench · Agentic Coding 60.9LiveBench · Mathematics 87.9LiveBench · Data Analysis 70.2LiveBench · Language 79.9LiveBench · Instruction Following 69.3
SimpleBench ↗1 measureread 2026-09-26
SimpleBench 66.2%
Vals.ai ↗13 measuresread 2026-09-26
Vals · Legal Research Bench 49.04%Vals · LegalBench 84.84%Vals · Harvey Legal Agent Benchmark 8.33%Vals · Finance Agent 55.84%Vals · TaxEval 72.36%Vals · MedCode 42.86%Vals · MedScribe 88.81%Vals · SWE-bench Verified 95.4% (Mini-SWE-agent)Vals · Vibe Code Bench 78.13% (OpenHands)Vals · Code Migration 44.22%Vals · GPQA Diamond 88.13%Vals · MMLU Pro 86.77%Vals · ProofBench 49%
GPQA Diamond (Pass@1) 88.1%Terminal-Bench 2.1 (Pass@1) 88.2%Terminal-Bench 3.0 (Pass@1) 28.3%Terminal-Bench 4.0 (Pass@1) 37.9%DeepSWE v1.1 (Resolved) 66.9%ProgramBench (Almost@1) 19%NL2Repo-Bench (Score) 58%CyberGym (Pass@1) 84.5%ExploitGym (Pass@1) 15%HLE w/ tools (Pass@1) 62.5%AutomationBench (Pass@1) 48.8%Agent's Last Exam (Pass@1) 28.5%

Badge

PublicAI Index badge for GLM-5.3[![PublicAI Index](https://publicai.io/model-index/badge?model=glm-5-3)](https://publicai.io/model-index/m/glm-5-3)

✱ Placed by a one-off publication, not a board that re-ran the model. ✱✱ marks a figure the model’s own publisher printed, which counts for less again. Scores are 0–100 on the PublicAI Index scale, 50 = the average of the models each source lists. Scores belong to their publishers.

Snapshot 2026-09-26 · the full index · JSON API