‹ PublicAI Index
The LLM benchmark aggregator.
Qwen3 32B
Alibaba · 33B · open weights
Strongest in Factual grounding (#23 of 101), weakest in Human preference (#169 of 342). Above par in 5 of 11 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 5 and GPT-4.1 and ahead of GLM-4.6 and Grok 4.3.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Coding49−20.1#81/1651/5
Agentic coding48.8−19.9#77/1571/4
Safety53.2−8.3#93/3371/3
Factual grounding60.3−10.3#23/1011/1
Core abilities48.6−18.8#112/2041/3
General intelligence48−22.1#116/2041/3
Agents47.5−20.6#145/2682/5
Tool use57.9−16.1#23/811/1
Knowledge work38.1−35.5#140/1781/2
Human preference51.9−15.6#169/3421/1
Human preference51.9−15.6#169/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Qwen3 32B, left for the other.
§ 3 · Sources
Where the numbers come from
6 publications, 6 figures. Every one links to the page it was read from.
LMArena Text 1347
GDPval-AA 19
Aider polyglot 40%
BFCL v4 48.71%
Kagi LLM Benchmark 48.7%
Vectara · Factual consistency 94.1%
Badge
[](https://publicai.io/model-index/m/qwen3-32b)