‹ PublicAI Index
The LLM benchmark aggregator.
Qwen3.6 27B
Alibaba · 28B · open weights
Strongest in Data analysis (#43 of 57), weakest in Core abilities (#204 of 204). Above par in 4 of 16 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Mistral Medium 3.5 and Grok 4.3.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Professional51−10.6#71/1681/1
Finance54.8−8.1#47/1511/1
Legal44−23.1#109/1511/1
Agents51.8−16.3#99/2682/5
Knowledge work52.3−21.3#65/1782/2
Coding40.3−28.8#151/1652/5
Code generation38.5−31.5#63/771/2
Agentic coding41−27.7#136/1572/4
Reasoning35.2−31.7#171/1781/4
Mathematics36.4−29.3#124/1401/2
Reasoning31.8−37.3#139/1401/3
Core abilities32.8−34.6#204/2041/3
Data analysis44.9−21.2#43/571/1
Language26−45.5#56/571/1
Instruction following26−47.7#56/571/1
General intelligence34−36.1#197/2041/3
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Qwen3.6 27B, left for the other.
§ 3 · Sources
Where the numbers come from
4 publications, 16 figures. Every one links to the page it was read from.
GDPval-AA 973
AA-Briefcase 813
LiveBench 64LiveBench · Reasoning 70.3LiveBench · Coding 71.8LiveBench · Agentic Coding 39.3LiveBench · Mathematics 79.9LiveBench · Data Analysis 70.4LiveBench · Language 63.3LiveBench · Instruction Following 53.2
Vals · CaseLaw 53.16%Vals · CorpFin 62.31%Vals · TaxEval 71.26%Vals · MortgageTax 68.28%Vals · SWE-bench Verified 70% (Mini-SWE-agent)Vals · Vibe Code Bench 11.94% (OpenHands)
Badge
[](https://publicai.io/model-index/m/qwen3-6-27b)