‹ PublicAI Index
The LLM benchmark aggregator.
Qwen3.8 Max
Alibaba
Strongest in Knowledge work (#5 of 178), weakest in Toxicity avoidance (#185 of 272). Above par in 26 of 28 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5.5 and Claude Fable 5 and ahead of GPT-6 Sol and Gemini 3.8 Flash.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents65−3.1#6/2682/5
Knowledge work69.5−4.1#5/1782/2
Core abilities58.7−8.7#13/2042/3
Instruction following60.9−12.8#12/571/1
General intelligence61.6−8.5#15/2042/3
Data analysis58.4−7.7#17/571/1
Language50.8−20.7#29/571/1
Knowledge57.7−5.8#21/1381/2
Academic knowledge59.2−4.7#18/1231/1
Human preference64.6−2.9#22/3421/1
Human preference64.6−2.9#22/3421/1
Professional56.1−5.5#27/1681/1
Legal59.8−7.3#14/1511/1
Finance56.3−6.6#32/1511/1
Medical52.4−14.4#68/1401/1
Reasoning55.9−11#31/1782/4
Science62.3−1.2#6/1221/1
Reasoning56−13.1#37/1401/3
Mathematics52.8−12.9#67/1402/2
Coding53.5−15.6#42/1652/5
Agentic coding57.1−11.6#24/1572/4
Code generation40.8−29.2#59/771/2
Safety53.1−8.4#100/3371/3
Secure code63.9−2.6#20/2741/1
Harm refusal52.4−9.2#120/3001/2
Fairness48.3−21.7#136/3001/2
Jailbreak resistance53.2−11.7#142/2721/1
Toxicity avoidance51.1−7.9#185/2721/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Qwen3.8 Max, left for the other.
§ 3 · Sources
Where the numbers come from
7 publications, 33 figures. Every one links to the page it was read from.
LMArena Text 1479
Artificial Analysis Intelligence Index 45
GDPval-AA 1668
AA-Briefcase 1626
LiveBench 78.5LiveBench · Reasoning 88.2LiveBench · Coding 72.9LiveBench · Agentic Coding 64.6LiveBench · Mathematics 91.3LiveBench · Data Analysis 78.4LiveBench · Language 79.7LiveBench · Instruction Following 74.1
Vals · Legal Research Bench 47.6%Vals · LegalBench 83.61%Vals · Harvey Legal Agent Benchmark 10.42%Vals · Finance Agent 50.59%Vals · CorpFin 65.85%Vals · TaxEval 75.55%Vals · MortgageTax 63.99%Vals · MedCode 40.67%Vals · MedScribe 84.95%Vals · SWE-bench Verified 85.6% (Mini-SWE-agent)Vals · Vibe Code Bench 64.7% (OpenHands)Vals · Code Migration 23.96%Vals · GPQA Diamond 93.69%Vals · MMLU Pro 88.6%Vals · ProofBench 58%
Enkrypt · Jailbreak risk 10.1%Enkrypt · Harmful content risk 0.6%Enkrypt · CBRN risk 26%Enkrypt · Toxicity risk 5.5%Enkrypt · Bias risk 80.4%Enkrypt · Insecure code risk 5.3%
Badge
[](https://publicai.io/model-index/m/qwen3-8-max)