‹ PublicAI Index
The LLM benchmark aggregator.
Qwen2 72B Instruct
Alibaba · 72B
Strongest in Toxicity avoidance (#28 of 272), weakest in Human preference (#249 of 342). Above par in 7 of 9 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 5 and GPT-5 and ahead of Gemini 2.5 Pro and DeepSeek R1.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety55.3−6.2#46/3372/3
Toxicity avoidance58.1−0.9#28/2721/1
Safe-prompt compliance56.2−4.2#28/821/1
Jailbreak resistance61.3−3.6#30/2721/1
Harm refusal55.2−6.4#72/3002/2
Secure code56.7−9.8#102/2741/1
Fairness50.8−19.2#103/3002/2
Human preference43.7−23.8#249/3421/1
Human preference43.7−23.8#249/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Qwen2 72B Instruct, left for the other.
§ 3 · Sources
Where the numbers come from
3 publications, 12 figures. Every one links to the page it was read from.
LMArena Text 1262
HELM Safety · HarmBench 76.8%HELM Safety · SimpleSafetyTests 98.5%HELM Safety · Anthropic Red Team 99.1%HELM Safety · BBQ 95.1%HELM Safety · XSTest 96.9%
Enkrypt · Jailbreak risk 3.2%Enkrypt · Harmful content risk 12.8%Enkrypt · CBRN risk 9.7%Enkrypt · Toxicity risk 0.6%Enkrypt · Bias risk 84.2%Enkrypt · Insecure code risk 20%
Badge
[](https://publicai.io/model-index/m/qwen2-72b-instruct)