‹ PublicAI Index
The LLM benchmark aggregator.
Qwen3.7 Plus
Alibaba
Strongest in Computer use (#6 of 37), weakest in Legal (#127 of 151). Above par in 5 of 12 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Opus 5 and ahead of GPT-5 Mini and GPT-4.1 Mini.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Human preference62.4−5.1#54/3421/1
Human preference62.4−5.1#54/3421/1
Agents55.1−13#63/2683/5
Computer use64.8−6.9#6/371/1
Knowledge work51.3−22.3#73/1782/2
Coding47.1−22#96/1651/5
Agentic coding46.7−22#90/1571/4
Core abilities49.7−17.7#102/2041/3
General intelligence49.6−20.5#104/2041/3
Professional46.1−15.5#120/1681/1
Finance48.7−14.2#94/1511/1
Legal41.6−25.5#127/1511/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Qwen3.7 Plus, left for the other.
§ 3 · Sources
Where the numbers come from
6 publications, 11 figures. Every one links to the page it was read from.
LMArena Text 1456
Artificial Analysis Intelligence Index 25
GDPval-AA 756
AA-Briefcase 912
OSWorld 73.3%
Vals · Legal Research Bench 16.35%Vals · Harvey Legal Agent Benchmark 0%Vals · Finance Agent 38.22%Vals · MortgageTax 66.18%Vals · Vibe Code Bench 46.39% (OpenHands)Vals · Code Migration 12.86%
Badge
[](https://publicai.io/model-index/m/qwen3-7-plus)