‹ PublicAI Index
The LLM benchmark aggregator.
Qwen1.5 72B Chat
Alibaba · 72B
Strongest in Safe-prompt compliance (#42 of 82), weakest in Human preference (#261 of 341). Above par in 3 of 6 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 4.5 and GPT-5.1 and ahead of GLM-5.2 and Mimo V2.6 Flash.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety50.3−11.2#176/3361/3
Safe-prompt compliance53.5−6.9#42/821/1
Harm refusal50.8−10.8#143/2991/2
Fairness46.2−23.8#169/2991/2
Human preference41−26.5#261/3411/1
Human preference41−26.5#261/3411/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Qwen1.5 72B Chat, left for the other.
§ 3 · Sources
Where the numbers come from
2 publications, 6 figures. Every one links to the page it was read from.
LMArena Text 1234
HELM Safety · HarmBench 64.8%HELM Safety · SimpleSafetyTests 99%HELM Safety · Anthropic Red Team 99%HELM Safety · BBQ 84.6%HELM Safety · XSTest 95.7%
Badge
[](https://publicai.io/model-index/m/qwen1-5-72b-chat)