‹ PublicAI Index
The LLM benchmark aggregator.
Qwen3.8 27B
Alibaba · 28B · open weights
Strongest in Instruction following (#16 of 57), weakest in Medical (#113 of 140). Above par in 19 of 26 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of GPT-6 Luna and O4 Mini.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents60.8−7.3#23/2682/5
Knowledge work64−9.6#20/1782/2
Tool use68.5−5.5—/81✱0/1
Coding53−16.1#48/1652/5
Agentic coding54.8−13.9#39/1572/4
Code generation46.5−23.5#55/771/2
Professional52−9.6#59/1681/1
Legal57.1−10#27/1511/1
Finance52.7−10.2#70/1511/1
Medical44.5−22.3#113/1401/1
Core abilities52.5−14.9#66/2042/3
Instruction following58.4−15.3#16/571/1
Data analysis55.4−10.7#24/571/1
Language40.6−30.9#44/571/1
General intelligence53.9−16.2#67/2042/3
Long context60.9−3.7—/0✱✱
Knowledge53.2−10.3#67/1381/2
Academic knowledge53.9−10#62/1231/1
Factuality49−10—/0✱✱
Human preference60.7−6.8#79/3421/1
Human preference60.7−6.8#79/3421/1
Reasoning47.6−19.3#103/1783/4
Science58.9−4.6#32/1221/1
Reasoning49.5−19.6#63/1402/3
Mathematics40.6−25.1#109/1402/2
Expert reasoning53.8−7.1—/0✱✱
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Qwen3.8 27B, left for the other.
§ 3 · Sources
Where the numbers come from
8 publications, 36 figures. Every one links to the page it was read from.
LMArena Text 1438
Artificial Analysis Intelligence Index 34
GDPval-AA 1409
AA-Briefcase 1400
LiveBench 75.3LiveBench · Reasoning 80LiveBench · Coding 75.7LiveBench · Agentic Coding 61.4LiveBench · Mathematics 86.2LiveBench · Data Analysis 76.6LiveBench · Language 74.3LiveBench · Instruction Following 72.7
SimpleBench 60.2%
Vals · Legal Research Bench 36.06%Vals · LegalBench 82.43%Vals · Harvey Legal Agent Benchmark 11.25%Vals · Finance Agent 48.55%Vals · TaxEval 70.85%Vals · MortgageTax 64.94%Vals · MedCode 28.7%Vals · MedScribe 83.85%Vals · SWE-bench Verified 86% (Mini-SWE-agent)Vals · Vibe Code Bench 64.85% (OpenHands)Vals · Code Migration 14.16%Vals · GPQA Diamond 88.89%Vals · MMLU Pro 84.34%Vals · ProofBench 16%
tau3-Banking 48%Terminal-Bench 2.1 79.8%SciCode 44.7%Humanity's Last Exam (without tools) 33.9%GPQA Diamond 90.5%CritPt 5.4%AA-LCR 77.3%AA-Omniscience Accuracy 15.6%AA-Omniscience Non-Hallucination 69.7%
Badge
[](https://publicai.io/model-index/m/qwen3-8-27b)