‹ PublicAI Index
The LLM benchmark aggregator.
Qwen 2.5 Instruct Turbo (7B)
Alibaba · 7B
Strongest in Safe-prompt compliance (#32 of 82), weakest in Harm refusal (#195 of 299). Above par in 4 of 9 scopes. Among the models it meets almost everywhere, it finishes behind GPT-5 and Claude Opus 4 and ahead of Grok 4 and DeepSeek V3.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Professional47.1−14.5#113/1681/1
Legal43.4−23.7#114/1511/1
Safety51.2−10.3#153/3362/3
Safe-prompt compliance55.5−4.9#32/821/1
Fairness61.4−8.6#34/2992/2
Toxicity avoidance52.7−6.3#162/2721/1
Jailbreak resistance48.3−16.6#190/2721/1
Secure code45.5−21#193/2741/1
Harm refusal47.8−13.8#195/2992/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Qwen 2.5 Instruct Turbo (7B), left for the other.
§ 3 · Sources
Where the numbers come from
3 publications, 12 figures. Every one links to the page it was read from.
Vals · LegalBench 69.56%
HELM Safety · HarmBench 67.7%HELM Safety · SimpleSafetyTests 96%HELM Safety · Anthropic Red Team 98.5%HELM Safety · BBQ 90.6%HELM Safety · XSTest 96.6%
Enkrypt · Jailbreak risk 14.2%Enkrypt · Harmful content risk 51.7%Enkrypt · CBRN risk 15.7%Enkrypt · Toxicity risk 4.4%Enkrypt · Bias risk 50.7%Enkrypt · Insecure code risk 42.7%
Badge
[](https://publicai.io/model-index/m/qwen-2-5-instruct-turbo-7b)