‹ PublicAI Index
The LLM benchmark aggregator.
Qwen2 57B A14B Instruct
Alibaba · 57B · A14B
Strongest in Toxicity avoidance (#58 of 272), weakest in Harm refusal (#185 of 299). Above par in 4 of 6 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 5 and Grok 4.7 and ahead of Mistral Large and Gemma 3 27B It.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety51.6−9.9#141/3361/3
Toxicity avoidance56.6−2.4#58/2721/1
Secure code56.7−9.8#105/2741/1
Fairness50−20#116/2991/2
Jailbreak resistance53.3−11.6#141/2721/1
Harm refusal48.2−13.4#185/2991/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Qwen2 57B A14B Instruct, left for the other.
§ 3 · Sources
Where the numbers come from
1 publication, 6 figures. Every one links to the page it was read from.
Enkrypt · Jailbreak risk 10%Enkrypt · Harmful content risk 44.4%Enkrypt · CBRN risk 13.7%Enkrypt · Toxicity risk 1.7%Enkrypt · Bias risk 77.8%Enkrypt · Insecure code risk 20%
Badge
[](https://publicai.io/model-index/m/qwen2-57b-a14b-instruct)