‹ PublicAI Index
The LLM benchmark aggregator.
Claude Sonnet 4.5
Anthropic
Strongest in Tool use (#2 of 81), weakest in Secure code (#204 of 274). Above par in 23 of 28 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 4.8 and Claude Fable 5 and ahead of Grok 4 and Qwen3.8 27B.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge56.4−7.1#31/1381/2
Academic knowledge57.7−6.2#28/1231/1
Safety56.5−5#32/3373/3
Fairness65.1−4.9#24/3002/2
Harm refusal57.6−4#26/3002/2
Safe-prompt compliance55.1−5.3#34/821/1
Jailbreak resistance60−4.9#43/2721/1
Toxicity avoidance57.1−1.9#47/2721/1
Factual grounding44.9−25.7#76/1011/1
Secure code42.7−23.8#204/2741/1
Professional54.7−6.9#36/1681/1
Medical55.8−11#30/1401/1
Legal55.2−11.9#39/1511/1
Finance54.1−8.8#58/1511/1
Agents58−10.1#44/2684/5
Tool use74leads#2/811/1
Computer use59.2−12.5#10/371/1
Knowledge work50.2−23.4#80/1782/2
Human preference62.5−5#49/3421/1
Human preference62.5−5#49/3421/1
Reasoning52.4−14.5#73/1783/4
Mathematics55.8−9.9#40/1401/2
Science53.9−9.6#63/1221/1
Reasoning48.9−20.2#66/1402/3
Core abilities51.8−15.6#73/2041/3
General intelligence52.5−17.6#74/2041/3
Coding46.5−22.6#104/1651/5
Agentic coding46−22.7#98/1571/4
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Claude Sonnet 4.5, left for the other.
§ 3 · Sources
Where the numbers come from
12 publications, 33 figures. Every one links to the page it was read from.
LMArena Text 1457
GDPval-AA 893
AA-Briefcase 713
ARC-AGI-2 13.6%
BFCL v4 73.24%
OSWorld 62.9%
Kagi LLM Benchmark 57.9%
SimpleBench 54.3%
Vals · CaseLaw 62.16%Vals · LegalBench 84.08%Vals · CorpFin 61.97%Vals · TaxEval 73.3%Vals · MortgageTax 63.99%Vals · MedQA 94.71%Vals · MedCode 44.13%Vals · MedScribe 84.1%Vals · SWE-bench Verified 70% (Mini-SWE-agent)Vals · Vibe Code Bench 22.62% (OpenHands)Vals · GPQA Diamond 81.63%Vals · MMLU Pro 87.36%Vals · AIME 88.19%
HELM Safety · HarmBench 92%HELM Safety · SimpleSafetyTests 100%HELM Safety · Anthropic Red Team 98.2%HELM Safety · BBQ 98.9%HELM Safety · XSTest 96.4%
Enkrypt · Jailbreak risk 4.3%Enkrypt · Harmful content risk 0.6%Enkrypt · CBRN risk 13%Enkrypt · Toxicity risk 1.3%Enkrypt · Bias risk 52.5%Enkrypt · Insecure code risk 48.4%
Vectara · Factual consistency 88%
Badge
[](https://publicai.io/model-index/m/claude-sonnet-4-5)