‹ PublicAI Index
The LLM benchmark aggregator.
Qwen3.7 Max
Alibaba
Strongest in Academic knowledge (#12 of 123), weakest in Harm refusal (#246 of 300). Above par in 17 of 28 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Gemini 3 Pro and ahead of Claude Sonnet 4 and GPT-4.1.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge58.4−5.1#14/1381/2
Academic knowledge60.1−3.8#12/1231/1
Human preference64.2−3.3#30/3421/1
Human preference64.2−3.3#30/3421/1
Reasoning52.8−14.1#68/1783/4
Science59.8−3.7#25/1221/1
Reasoning55.2−13.9#41/1402/3
Mathematics44.1−21.6#100/1401/2
Professional51.2−10.4#70/1681/1
Finance54.8−8.1#48/1511/1
Legal49.6−17.5#82/1511/1
Medical48.6−18.2#91/1401/1
Agents53.9−14.2#77/2682/5
Knowledge work55−18.6#54/1782/2
Core abilities51.3−16.1#79/2041/3
Instruction following60.7−13#13/571/1
Language50.8−20.7#31/571/1
Data analysis47.3−18.8#39/571/1
General intelligence47.3−22.8#121/2041/3
Coding44.3−24.8#122/1652/5
Code generation43.4−26.6#57/771/2
Agentic coding44.5−24.2#108/1572/4
Safety48.3−13.2#225/3371/3
Secure code62.1−4.4#47/2741/1
Fairness52.2−17.8#94/3001/2
Toxicity avoidance52.7−6.3#157/2721/1
Jailbreak resistance33.3−31.6#241/2721/1
Harm refusal45−16.6#246/3001/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Qwen3.7 Max, left for the other.
§ 3 · Sources
Where the numbers come from
7 publications, 31 figures. Every one links to the page it was read from.
LMArena Text 1475
GDPval-AA 1114
AA-Briefcase 915
LiveBench 73.1LiveBench · Reasoning 83.3LiveBench · Coding 74.2LiveBench · Agentic Coding 43.6LiveBench · Mathematics 85.2LiveBench · Data Analysis 71.8LiveBench · Language 79.7LiveBench · Instruction Following 74
SimpleBench 70.4%
Vals · Legal Research Bench 25.48%Vals · LegalBench 84.91%Vals · Harvey Legal Agent Benchmark 1.67%Vals · Finance Agent 47.78%Vals · CorpFin 63.71%Vals · TaxEval 75.31%Vals · MedCode 38.75%Vals · MedScribe 79.4%Vals · SWE-bench Verified 68.8% (Mini-SWE-agent)Vals · Vibe Code Bench 47.67% (OpenHands)Vals · Code Migration 13.07%Vals · GPQA Diamond 90.15%Vals · MMLU Pro 89.31%
Enkrypt · Jailbreak risk 26.9%Enkrypt · Harmful content risk 8.9%Enkrypt · CBRN risk 49.8%Enkrypt · Toxicity risk 4.4%Enkrypt · Bias risk 74.4%Enkrypt · Insecure code risk 8.9%
Badge
[](https://publicai.io/model-index/m/qwen3-7-max)