‹ PublicAI Index
The LLM benchmark aggregator.
Claude Sonnet 4
Anthropic
Strongest in Toxicity avoidance (#12 of 272), weakest in Agents (#127 of 268). Above par in 17 of 28 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Kimi K2.6 and Claude 3.5 Sonnet.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Core abilities56.9−10.5#27/2041/3
General intelligence59.9−10.2#28/2041/3
Safety56.6−4.9#28/3373/3
Toxicity avoidance58.7−0.3#12/2721/1
Jailbreak resistance61.8−3.1#23/2721/1
Fairness62.1−7.9#31/3002/2
Safe-prompt compliance55.5−4.9#31/821/1
Factual grounding49.2−21.4#58/1011/1
Secure code59.3−7.2#70/2741/1
Harm refusal54.2−7.4#84/3002/2
Coding52.7−16.4#54/1652/5
Agentic coding54.4−14.3#44/1571/4
Code generation50−20#46/771/2
Knowledge52.7−10.8#73/1381/2
Academic knowledge53.3−10.6#67/1231/1
Professional48.7−12.9#94/1681/1
Legal52.9−14.2#54/1511/1
Finance48.7−14.2#95/1511/1
Medical46.1−20.7#107/1401/1
Reasoning48.6−18.3#97/1783/4
Mathematics52.7−13#69/1401/2
Science49.2−14.3#79/1221/1
Reasoning45.1−24#82/1402/3
Human preference57.2−10.3#122/3421/1
Human preference57.2−10.3#122/3421/1
Agents49.3−18.8#127/2682/5
Computer use48.9−22.8#18/371/1
Knowledge work49.4−24.2#87/1781/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Claude Sonnet 4, left for the other.
§ 3 · Sources
Where the numbers come from
12 publications, 30 figures. Every one links to the page it was read from.
LMArena Text 1402
GDPval-AA 680
ARC-AGI-2 5.9%
LiveCodeBench 55.9%
Aider polyglot 61.3%
OSWorld 43.9%
Kagi LLM Benchmark 73%
SimpleBench 45.5%
Vals · LegalBench 82.06%Vals · CorpFin 61.23%Vals · TaxEval 72%Vals · MortgageTax 49.88%Vals · MedQA 92.71%Vals · MedCode 34.96%Vals · MedScribe 69.35%Vals · GPQA Diamond 75%Vals · MMLU Pro 83.86%Vals · AIME 76.25%
HELM Safety · HarmBench 98%HELM Safety · SimpleSafetyTests 100%HELM Safety · Anthropic Red Team 99.2%HELM Safety · BBQ 96.6%HELM Safety · XSTest 96.6%
Enkrypt · Jailbreak risk 2.8%Enkrypt · Harmful content risk 0.6%Enkrypt · CBRN risk 37%Enkrypt · Toxicity risk 0.2%Enkrypt · Bias risk 57.1%Enkrypt · Insecure code risk 14.7%
Vectara · Factual consistency 89.7%
Badge
[](https://publicai.io/model-index/m/claude-sonnet-4)