‹ PublicAI Index
The LLM benchmark aggregator.
GLM-4.7
Z.ai · 358B · open weights
Strongest in Mathematics (#24 of 140, on 1 of its 2 boards), weakest in Finance (#121 of 151). Above par in 9 of 16 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Opus 4.7 and ahead of GPT-5.4 Mini and DeepSeek R1.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Reasoning53.1−13.8#62/1782/4
Mathematics57.2−8.5#24/1401/2
Reasoning48.9−20.2#67/1401/3
Science52.8−10.7#69/1221/1
Human preference61−6.5#75/3421/1
Human preference61−6.5#75/3421/1
Knowledge51.6−11.9#78/1381/2
Academic knowledge51.9−12#72/1231/1
Agents53.2−14.9#85/2681/5
Knowledge work54.8−18.8#55/1781/2
Coding48.5−20.6#86/1651/5
Agentic coding48.2−20.5#82/1571/4
Professional46.2−15.4#119/1681/1
Legal49.8−17.3#78/1511/1
Medical45.1−21.7#112/1401/1
Finance43.1−19.8#121/1511/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GLM-4.7, left for the other.
§ 3 · Sources
Where the numbers come from
4 publications, 14 figures. Every one links to the page it was read from.
LMArena Text 1441
GDPval-AA 1000
SimpleBench 47.7%
Vals · CaseLaw 54.88%Vals · LegalBench 83.36%Vals · CorpFin 46.39%Vals · TaxEval 68.77%Vals · MedQA 93.74%Vals · MedCode 32.77%Vals · MedScribe 68.63%Vals · SWE-bench Verified 69.4% (Mini-SWE-agent)Vals · GPQA Diamond 80.05%Vals · MMLU Pro 82.74%Vals · AIME 93.33%
Badge
[](https://publicai.io/model-index/m/glm-4-7)