‹ PublicAI Index
The LLM benchmark aggregator.
GLM-4.6
Z.ai · 357B · open weights
Strongest in Tool use (#4 of 81), weakest in Harm refusal (#264 of 300). Above par in 15 of 25 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5 and Claude Opus 5.5 and ahead of GLM-4.5 and Gemini 3.5 Flash Lite.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents59.5−8.6#36/2682/5
Tool use74leads#4/811/1
Knowledge work50.5−23.1#78/1781/2
Reasoning53−13.9#64/1781/4
Mathematics57−8.7#27/1401/2
Science48.9−14.6#80/1221/1
Professional50.4−11.2#76/1681/1
Medical54.7−12.1#44/1401/1
Legal51.1−16#69/1511/1
Finance47.5−15.4#102/1511/1
Knowledge51−12.5#81/1381/2
Academic knowledge51.2−12.7#75/1231/1
Human preference59.3−8.2#98/3421/1
Human preference59.3−8.2#98/3421/1
Core abilities47.6−19.8#119/2041/3
General intelligence46.6−23.5#125/2041/3
Coding43.3−25.8#129/1651/5
Agentic coding41.9−26.8#130/1571/4
Safety47.1−14.4#254/3372/3
Factual grounding51.2−19.4#51/1011/1
Secure code56−10.5#110/2741/1
Toxicity avoidance52.4−6.6#168/2721/1
Jailbreak resistance40.6−24.3#216/2721/1
Fairness43.9−26.1#223/3001/2
Harm refusal43.6−18#264/3001/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GLM-4.6, left for the other.
§ 3 · Sources
Where the numbers come from
7 publications, 19 figures. Every one links to the page it was read from.
LMArena Text 1424
GDPval-AA 746
BFCL v4 72.38%
Kagi LLM Benchmark 45.7%
Vals · LegalBench 79.61%Vals · CorpFin 56.84%Vals · TaxEval 66.23%Vals · MedQA 92.22%Vals · Vibe Code Bench 3.09% (OpenHands)Vals · GPQA Diamond 74.5%Vals · MMLU Pro 82.2%Vals · AIME 92.71%
Enkrypt · Jailbreak risk 20.7%Enkrypt · Harmful content risk 45%Enkrypt · CBRN risk 25.3%Enkrypt · Toxicity risk 4.6%Enkrypt · Bias risk 87.3%Enkrypt · Insecure code risk 21.3%
Vectara · Factual consistency 90.5%
Badge
[](https://publicai.io/model-index/m/glm-4-6)