‹ PublicAI Index
The LLM benchmark aggregator.
Qwen3.5 35B A3B
Alibaba · 36B · A3B · open weights
Strongest in Secure code (#40 of 274), weakest in Jailbreak resistance (#183 of 272). Above par in 7 of 11 scopes. Among the models it meets almost everywhere, it finishes behind Grok 4.7 and GPT-5.5 and ahead of Kimi K2.6 and Kimi K2.5.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety53.2−8.3#95/3372/3
Secure code62.6−3.9#40/2741/1
Toxicity avoidance56.7−2.3#53/2721/1
Factual grounding48.7−21.9#62/1011/1
Fairness52.9−17.1#84/3001/2
Harm refusal51.9−9.7#127/3001/2
Jailbreak resistance49.3−15.6#183/2721/1
Human preference56.4−11.1#130/3421/1
Human preference56.4−11.1#130/3421/1
Agents48.6−19.5#134/2681/5
Knowledge work47.9−25.7#92/1781/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Qwen3.5 35B A3B, left for the other.
§ 3 · Sources
Where the numbers come from
4 publications, 9 figures. Every one links to the page it was read from.
LMArena Text 1394
GDPval-AA 595
Enkrypt · Jailbreak risk 13.4%Enkrypt · Harmful content risk 1.7%Enkrypt · CBRN risk 26.8%Enkrypt · Toxicity risk 1.6%Enkrypt · Bias risk 73.4%Enkrypt · Insecure code risk 8%
Vectara · Factual consistency 89.5%
Badge
[](https://publicai.io/model-index/m/qwen3-5-35b-a3b)