‹ PublicAI Index
The LLM benchmark aggregator.
Qwen2.5 3B Instruct
Alibaba · 3B
Strongest in Secure code (#147 of 274), weakest in Fairness (#252 of 299). Above par in 2 of 6 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 5 and Claude Opus 5.5 and ahead of Open Codestral Mamba and Qwen3 8B.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety47.3−14.2#245/3361/3
Secure code51.4−15.1#147/2741/1
Toxicity avoidance52−7#179/2721/1
Jailbreak resistance47.6−17.3#194/2721/1
Harm refusal45.6−16#234/2991/2
Fairness42.6−27.4#252/2991/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Qwen2.5 3B Instruct, left for the other.
§ 3 · Sources
Where the numbers come from
1 publication, 6 figures. Every one links to the page it was read from.
Enkrypt · Jailbreak risk 14.8%Enkrypt · Harmful content risk 58.3%Enkrypt · CBRN risk 13%Enkrypt · Toxicity risk 4.9%Enkrypt · Bias risk 89.4%Enkrypt · Insecure code risk 30.7%
Badge
[](https://publicai.io/model-index/m/qwen2-5-3b-instruct)