‹ PublicAI Index
The LLM benchmark aggregator.
Qwen3.5 Flash
Alibaba
Strongest in Mathematics (#29 of 140, on 1 of its 2 boards), weakest in Safety (#193 of 337). Above par in 9 of 15 scopes. Among the models it meets almost everywhere, it finishes behind GPT-5.5 and GPT-5.4 and ahead of Kimi K2.5 and Gemini 3.1 Flash Lite.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Reasoning55−11.9#42/1781/4
Mathematics57−8.7#29/1401/2
Science54.7−8.8#60/1221/1
Knowledge53−10.5#72/1381/2
Academic knowledge53.5−10.4#66/1231/1
Professional50−11.6#79/1681/1
Finance55.4−7.5#42/1511/1
Legal51−16.1#71/1511/1
Medical40.9−25.9#124/1401/1
Coding46.7−22.4#101/1651/5
Agentic coding46−22.7#97/1571/4
Human preference56.6−10.9#126/3421/1
Human preference56.6−10.9#126/3421/1
Safety49.6−11.9#193/3371/3
Factual grounding48.7−21.9#60/1011/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Qwen3.5 Flash, left for the other.
§ 3 · Sources
Where the numbers come from
3 publications, 13 figures. Every one links to the page it was read from.
LMArena Text 1396
Vals · CaseLaw 55.95%Vals · LegalBench 84.28%Vals · CorpFin 63.56%Vals · TaxEval 72.16%Vals · MortgageTax 67.37%Vals · MedCode 33%Vals · MedScribe 70.62%Vals · SWE-bench Verified 64.4% (Mini-SWE-agent)Vals · GPQA Diamond 82.83%Vals · MMLU Pro 84.06%Vals · AIME 92.5%
Vectara · Factual consistency 89.5%
Badge
[](https://publicai.io/model-index/m/qwen3-5-flash)