‹ PublicAI Index
The LLM benchmark aggregator.
Qwen3.5 2B
Alibaba · 2B
Strongest in Knowledge work (#150 of 177, on 1 of its 2 boards), weakest in Agents (#232 of 267). Above par in 3 of 8 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 5 and GLM-5.2 and ahead of Kimi K2.6 and GPT-5.4 Mini.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents41.1−27#232/2671/5
Knowledge work36.7−36.9#150/1771/2
Tool use52−22—/81✱0/1
Reasoning49.2−17.7—/178✱0/4
Mathematics49.3−16.4—/140✱0/2
Science43.4−20.1—/122✱0/1
Coding51.4−17.7—/164✱0/5
Code generation53.6−16.4—/77✱0/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Qwen3.5 2B, left for the other.
§ 3 · Sources
Where the numbers come from
2 publications, 9 figures. Every one links to the page it was read from.
GDPval-AA -61
GPQA Diamond 54.9%HMMT Feb 2026 22.7%BFCL v4 43.6%HumanEval+ 75.6%MBPP+ 67.7%AIME 2025 34.2%AIME 2026 38.8%LiveCodeBench v6 29.8%
Badge
[](https://publicai.io/model-index/m/qwen3-5-2b)