‹ PublicAI Index
The LLM benchmark aggregator.
Qwen3.5 0.8B
Alibaba · 0.8B
Strongest in Knowledge work (#172 of 177, on 1 of its 2 boards), weakest in Agents (#262 of 267). Above par in 0 of 8 scopes. Among the models it meets almost everywhere, it finishes behind Qwen3.8 27B and Claude Sonnet 5 and ahead of Mistral Large and Mercury 2.5.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents37.2−30.9#262/2671/5
Knowledge work30.8−42.8#172/1771/2
Tool use36.6−37.4—/81✱0/1
Reasoning44.6−22.3—/178✱0/4
Mathematics40.6−25.1—/140✱0/2
Science31.5−32—/122✱0/1
Coding45.6−23.5—/164✱0/5
Code generation38.7−31.3—/77✱0/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Qwen3.5 0.8B, left for the other.
§ 3 · Sources
Where the numbers come from
2 publications, 9 figures. Every one links to the page it was read from.
GDPval-AA -405
GPQA Diamond 11.9%HMMT Feb 2026 0.6%BFCL v4 25.3%HumanEval+ 16.5%MBPP+ 35.4%AIME 2025 1%AIME 2026 0.2%LiveCodeBench v6 6.6%
Badge
[](https://publicai.io/model-index/m/qwen3-5-0-8b)