‹ PublicAI Index
The LLM benchmark aggregator.
QwQ 32B
Alibaba · 32B
Strongest in Agentic coding (#115 of 157, on 1 of its 4 boards), weakest in Fairness (#277 of 300). Above par in 4 of 10 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5.5 and Claude Opus 4.7 and ahead of Grok 4.1 Fast and GLM-4.6.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Coding44.9−24.2#116/1651/5
Agentic coding43.9−24.8#115/1571/4
Human preference50.8−16.7#183/3421/1
Human preference50.8−16.7#183/3421/1
Safety47.9−13.6#236/3371/3
Jailbreak resistance54.6−10.3#122/2721/1
Toxicity avoidance54.6−4.4#123/2721/1
Secure code47−19.5#185/2741/1
Harm refusal46.3−15.3#218/3001/2
Fairness40.2−29.8#277/3001/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for QwQ 32B, left for the other.
§ 3 · Sources
Where the numbers come from
3 publications, 8 figures. Every one links to the page it was read from.
LMArena Text 1336
Aider polyglot 20.9%
Enkrypt · Jailbreak risk 8.9%Enkrypt · Harmful content risk 66.1%Enkrypt · CBRN risk 7%Enkrypt · Toxicity risk 3.1%Enkrypt · Bias risk 93%Enkrypt · Insecure code risk 39.6%
Badge
[](https://publicai.io/model-index/m/qwq-32b)