‹ PublicAI Index
The LLM benchmark aggregator.
Qwen3.6 Plus
Alibaba
Strongest in Academic knowledge (#24 of 123), weakest in Core abilities (#187 of 204). Above par in 9 of 22 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Opus 4.7 and ahead of Claude Sonnet 4 and Claude Opus 4.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge56.7−6.8#27/1381/2
Academic knowledge58.1−5.8#24/1231/1
Human preference61.2−6.3#70/3421/1
Human preference61.2−6.3#70/3421/1
Agents52.9−15.2#87/2681/5
Knowledge work54.4−19.2#58/1781/2
Reasoning48.3−18.6#99/1782/4
Science57.9−5.6#38/1221/1
Mathematics49.7−16#80/1402/2
Reasoning39−30.1#116/1401/3
Professional48.2−13.4#101/1681/1
Finance53.2−9.7#67/1511/1
Medical46.3−20.5#104/1401/1
Legal44−23.1#108/1511/1
Coding44.5−24.6#118/1652/5
Code generation51.6−18.4#41/771/2
Agentic coding42.4−26.3#126/1572/4
Core abilities40−27.4#187/2041/3
Language41.9−29.6#42/571/1
Data analysis44.1−22#46/571/1
Instruction following33.1−40.6#52/571/1
General intelligence40.5−29.6#171/2041/3
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Qwen3.6 Plus, left for the other.
§ 3 · Sources
Where the numbers come from
4 publications, 26 figures. Every one links to the page it was read from.
LMArena Text 1443
GDPval-AA 976
LiveBench 68.9LiveBench · Reasoning 75.8LiveBench · Coding 78.2LiveBench · Agentic Coding 41.4LiveBench · Mathematics 83.7LiveBench · Data Analysis 69.9LiveBench · Language 75LiveBench · Instruction Following 58.3
Vals · Legal Research Bench 14.9%Vals · CaseLaw 51.45%Vals · LegalBench 84.23%Vals · Harvey Legal Agent Benchmark 1.25%Vals · Finance Agent 40.85%Vals · CorpFin 61.93%Vals · TaxEval 74.73%Vals · MortgageTax 67.97%Vals · MedCode 36.89%Vals · MedScribe 76.96%Vals · SWE-bench Verified 73.4% (Mini-SWE-agent)Vals · Vibe Code Bench 25.57% (OpenHands)Vals · Code Migration 11.1%Vals · GPQA Diamond 87.37%Vals · MMLU Pro 87.67%Vals · AIME 94.58%
Badge
[](https://publicai.io/model-index/m/qwen3-6-plus)