‹ PublicAI Index
The LLM benchmark aggregator.
Qwen 2.5 Instruct Turbo (72B)
Alibaba · 72B
Strongest in Safe-prompt compliance (#13 of 82), weakest in Secure code (#141 of 274). Above par in 8 of 10 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 4.8 and Grok 4.7 and ahead of GLM-5.3 Flash and Qwen3.8 Max.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety56.5−5#34/3372/3
Safe-prompt compliance58.4−2#13/821/1
Fairness65.4−4.6#22/3002/2
Toxicity avoidance55.8−3.2#88/2721/1
Harm refusal53.9−7.7#91/3002/2
Jailbreak resistance56.4−8.5#95/2721/1
Secure code51.7−14.8#141/2741/1
Professional49.3−12.3#89/1681/1
Legal50.9−16.2#72/1511/1
Medical46.7−20.1#98/1401/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Qwen 2.5 Instruct Turbo (72B), left for the other.
§ 3 · Sources
Where the numbers come from
3 publications, 13 figures. Every one links to the page it was read from.
Vals · LegalBench 79.4%Vals · MedQA 77.39%
HELM Safety · HarmBench 72.8%HELM Safety · SimpleSafetyTests 100%HELM Safety · Anthropic Red Team 99.6%HELM Safety · BBQ 95.4%HELM Safety · XSTest 97.9%
Enkrypt · Jailbreak risk 7.4%Enkrypt · Harmful content risk 28.3%Enkrypt · CBRN risk 8.8%Enkrypt · Toxicity risk 2.2%Enkrypt · Bias risk 44.2%Enkrypt · Insecure code risk 30.2%
Badge
[](https://publicai.io/model-index/m/qwen-2-5-instruct-turbo-72b)