‹ PublicAI Index
The LLM benchmark aggregator.
Qwen Max
Alibaba
Strongest in Toxicity avoidance (#67 of 272), weakest in Human preference (#204 of 342). Above par in 5 of 10 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 4.7 and Claude Sonnet 5 and ahead of DeepSeek V4 Pro and Mimo V2.6 Flash.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety53.3−8.2#91/3371/3
Toxicity avoidance56.3−2.7#67/2721/1
Jailbreak resistance57.7−7.2#73/2721/1
Secure code56.9−9.6#98/2741/1
Harm refusal52.8−8.8#113/3001/2
Fairness46.2−23.8#167/3001/2
Coding45.1−24#114/1651/5
Agentic coding44.1−24.6#112/1571/4
Human preference49.1−18.4#204/3421/1
Human preference49.1−18.4#204/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Qwen Max, left for the other.
§ 3 · Sources
Where the numbers come from
3 publications, 8 figures. Every one links to the page it was read from.
LMArena Text 1318
Aider polyglot 21.8%
Enkrypt · Jailbreak risk 6.3%Enkrypt · Harmful content risk 32.8%Enkrypt · CBRN risk 7.7%Enkrypt · Toxicity risk 1.9%Enkrypt · Bias risk 83.7%Enkrypt · Insecure code risk 19.6%
Badge
[](https://publicai.io/model-index/m/qwen-max)