§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents50.3−17.8—/267✱0/5
Tool use50.7−23.3—/81✱0/1
Reasoning50.7−16.2—/178✱0/4
Mathematics53.6−12.1—/140✱0/2
Science52.6−10.9—/122✱0/1
Expert reasoning48.6−12.3—/0✱✱
Coding47.5−21.6—/164✱0/5
Agentic coding48.4−20.3—/157✱0/4
Code generation37.1−32.9—/77✱0/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Qwen3.5 4B, left for the other.
§ 3 · Sources
Where the numbers come from
1 publication, 8 figures. Every one links to the page it was read from.
tau3-Banking 6.8%Terminal-Bench 2.1 25.8%SciCode 16.1%GPQA Diamond 77.1%SWE-bench Verified 41.2%HMMT Feb 2026 61.6%HLE 9.9%BFCL v4 55.7%
Badge
[](https://publicai.io/model-index/m/qwen3-5-4b)