‹ PublicAI Index
The LLM benchmark aggregator.
GLM-5
Z.ai · 754B · open weights
Strongest in Core abilities (#20 of 204, on 1 of its 3 boards), weakest in Safety (#183 of 337). Above par in 12 of 18 scopes. Among the models it meets almost everywhere, it finishes behind GPT-5.5 and Claude Fable 5 and ahead of O3 and O1.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Core abilities57.6−9.8#20/2041/3
General intelligence60.9−9.2#21/2041/3
Human preference62.5−5#50/3421/1
Human preference62.5−5#50/3421/1
Knowledge55−8.5#54/1381/2
Academic knowledge56−7.9#50/1231/1
Professional51.7−9.9#64/1681/1
Medical55.7−11.1#31/1401/1
Finance52.7−10.2#71/1511/1
Legal48.6−18.5#89/1511/1
Reasoning52.4−14.5#72/1783/4
Mathematics56.8−8.9#32/1401/2
Science55.1−8.4#58/1221/1
Reasoning47.5−21.6#71/1402/3
Coding47−22.1#99/1651/5
Agentic coding46.6−22.1#93/1571/4
Safety49.9−11.6#183/3371/3
Factual grounding49.7−20.9#57/1011/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GLM-5, left for the other.
§ 3 · Sources
Where the numbers come from
6 publications, 15 figures. Every one links to the page it was read from.
LMArena Text 1457
ARC-AGI-2 4.9%
Kagi LLM Benchmark 75%
SimpleBench 53.2%
Vals · CaseLaw 52.52%Vals · LegalBench 84.06%Vals · CorpFin 62.9%Vals · TaxEval 70.03%Vals · MedQA 94.27%Vals · SWE-bench Verified 71.4% (Mini-SWE-agent)Vals · Vibe Code Bench 23.36% (OpenHands)Vals · GPQA Diamond 83.33%Vals · MMLU Pro 86.03%Vals · AIME 91.67%
Vectara · Factual consistency 89.9%
Badge
[](https://publicai.io/model-index/m/glm-5)