‹ PublicAI Index
The LLM benchmark aggregator.
Qwen3 Max
Alibaba
Strongest in Finance (#24 of 151), weakest in Agentic coding (#129 of 157). Above par in 10 of 15 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Gemini 3.5 Flash Lite and Grok 4.3.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Core abilities56.7−10.7#30/2041/3
General intelligence59.7−10.4#30/2041/3
Reasoning54−12.9#47/1781/4
Science56.1−7.4#50/1221/1
Mathematics53.9−11.8#62/1401/2
Knowledge53.9−9.6#63/1381/2
Academic knowledge54.7−9.2#58/1231/1
Professional49.8−11.8#81/1681/1
Finance57.3−5.6#24/1511/1
Legal49.1−18#84/1511/1
Medical44.3−22.5#114/1401/1
Human preference59.2−8.3#100/3421/1
Human preference59.2−8.3#100/3421/1
Coding43.4−25.7#127/1651/5
Agentic coding42−26.7#129/1571/4
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Qwen3 Max, left for the other.
§ 3 · Sources
Where the numbers come from
3 publications, 13 figures. Every one links to the page it was read from.
LMArena Text 1423
Kagi LLM Benchmark 72.5%
Vals · CaseLaw 54.98%Vals · LegalBench 81.86%Vals · CorpFin 68.03%Vals · TaxEval 73.51%Vals · MedQA 87.37%Vals · MedCode 31.37%Vals · MedScribe 72.71%Vals · Vibe Code Bench 3.51% (OpenHands)Vals · GPQA Diamond 84.85%Vals · MMLU Pro 84.98%Vals · AIME 81.04%
Badge
[](https://publicai.io/model-index/m/qwen3-max)