‹ PublicAI Index
The LLM benchmark aggregator.
Claude Sonnet 5
Anthropic
Strongest in Safety (#1 of 337, on 1 of its 3 boards), weakest in Core abilities (#107 of 204). Above par in 30 of 34 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Grok 4.5 and Gemini 3.1 Pro.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety61.5leads#1/3371/3
Fairness70leads#5/3001/2
Toxicity avoidance58.7−0.3#11/2721/1
Jailbreak resistance62.8−2.1#12/2721/1
Harm refusal59.6−2#13/3001/2
Secure code62.1−4.4#45/2741/1
Reasoning57.7−9.2#24/1783/4
Mathematics57.7−8#21/1402/2
Reasoning57.1−12#31/1402/3
Science58.9−4.6#33/1221/1
Expert reasoning59.4−1.5—/0✱✱
Agents60.7−7.4#25/2682/5
Knowledge work63.9−9.7#22/1782/2
Tool use59.4−14.6—/81✱0/1
Web research58.3−5.9—/4✱0/1
Workflow automation59.3−2.7—/0✱✱
Coding56.5−12.6#27/1653/5
Code generation56.8−13.2#21/771/2
Agentic coding56.4−12.3#30/1573/4
Professional55.9−5.7#29/1681/1
Finance59.2−3.7#7/1511/1
Legal55.2−11.9#40/1511/1
Medical52.4−14.4#67/1401/1
Knowledge56.6−6.9#29/1381/2
Academic knowledge57.9−6#26/1231/1
Factuality59leads—/0✱✱
Human preference63−4.5#46/3421/1
Human preference63−4.5#46/3421/1
Core abilities48.9−18.5#107/2042/3
Data analysis47.1−19#40/571/1
Instruction following43−30.7#40/571/1
Language41.9−29.6#41/571/1
General intelligence56.3−13.8#54/2042/3
Long context60.6−4—/0✱✱
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Claude Sonnet 5, left for the other.
§ 3 · Sources
Where the numbers come from
10 publications, 50 figures. Every one links to the page it was read from.
LMArena Text 1462
Artificial Analysis Intelligence Index 38
GDPval-AA 1449
AA-Briefcase 1359
Terminal-Bench 12.4% (Claude Code)
LiveBench 76LiveBench · Reasoning 88.7LiveBench · Coding 80.7LiveBench · Agentic Coding 59.4LiveBench · Mathematics 92.9LiveBench · Data Analysis 71.7LiveBench · Language 75LiveBench · Instruction Following 63.9
SimpleBench 60.6%
Vals · Legal Research Bench 41.83%Vals · LegalBench 83.92%Vals · Harvey Legal Agent Benchmark 5%Vals · Finance Agent 53.91%Vals · CorpFin 66.98%Vals · TaxEval 75.63%Vals · MortgageTax 70.03%Vals · MedCode 47.54%Vals · MedScribe 76.05%Vals · SWE-bench Verified 79.6% (Mini-SWE-agent)Vals · Vibe Code Bench 81.33% (OpenHands)Vals · Code Migration 44.39%Vals · GPQA Diamond 88.89%Vals · MMLU Pro 87.55%Vals · ProofBench 77%
Enkrypt · Jailbreak risk 2%Enkrypt · Harmful content risk 0%Enkrypt · CBRN risk 7.5%Enkrypt · Toxicity risk 0.2%Enkrypt · Bias risk 29.5%Enkrypt · Insecure code risk 8.9%
GDPVal-AA 1584tau3-Banking 37.3%Toolathlon Verified 71.6%Automation Bench Public 34.7%Apex-Agents (pass@1) 31.7%MCPMark 65.3%BrowseComp 84.7%Terminal-Bench 2.1 80.5%SciCode 53.6%Humanity's Last Exam (without tools) 41.3%GPQA Diamond 91.1%CritPt 16.9%AA-LCR 77%AA-Omniscience Accuracy 40%AA-Omniscience Non-Hallucination 61%
Badge
[](https://publicai.io/model-index/m/claude-sonnet-5)