‹ PublicAI Index
The LLM benchmark aggregator.
Qwen 3.5 Plus
Alibaba
Strongest in Medical (#28 of 140), weakest in Safety (#201 of 337). Above par in 9 of 13 scopes. Among the models it meets almost everywhere, it finishes behind GPT-5.5 and Claude Opus 5 and ahead of Kimi K2.6 and Qwen3.5 Flash.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge56.2−7.3#36/1381/2
Academic knowledge57.4−6.5#33/1231/1
Reasoning55.2−11.7#40/1781/4
Science57.9−5.6#39/1221/1
Mathematics55.3−10.4#44/1401/2
Professional54−7.6#42/1681/1
Medical56.2−10.6#28/1401/1
Legal54−13.1#45/1511/1
Finance53.7−9.2#64/1511/1
Coding46−23.1#110/1651/5
Agentic coding45.4−23.3#103/1571/4
Safety49.4−12.1#201/3371/3
Factual grounding48.2−22.4#65/1011/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Qwen 3.5 Plus, left for the other.
§ 3 · Sources
Where the numbers come from
2 publications, 11 figures. Every one links to the page it was read from.
Vals · CaseLaw 59.7%Vals · LegalBench 85.1%Vals · CorpFin 65.31%Vals · MortgageTax 60.77%Vals · MedQA 95.21%Vals · SWE-bench Verified 71.2% (Mini-SWE-agent)Vals · Vibe Code Bench 15.74% (OpenHands)Vals · GPQA Diamond 87.37%Vals · MMLU Pro 87.18%Vals · AIME 86.04%
Vectara · Factual consistency 89.3%
Badge
[](https://publicai.io/model-index/m/qwen-3-5-plus)