‹ PublicAI Index
The LLM benchmark aggregator.
Claude 3.5 Sonnet
Anthropic
Strongest in Harm refusal (#1 of 300), weakest in Secure code (#155 of 274). Above par in 12 of 22 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of DeepSeek V3.2 and Kimi K2 Thinking.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety60.3−1.2#5/3372/3
Harm refusal61.6leads#1/3002/2
Jailbreak resistance64.9leads#1/2721/1
Toxicity avoidance58.8−0.2#8/2721/1
Fairness65.2−4.8#23/3002/2
Safe-prompt compliance53.3−7.1#43/821/1
Secure code50.3−16.2#155/2741/1
Knowledge59.1−4.4#9/1382/2
Multimodal understanding71.6−0.5#2/191/1
Academic knowledge46.5−17.4#95/1231/1
Professional49.7−11.9#83/1681/1
Medical49.8−17#81/1401/1
Finance49.5−13.4#91/1511/1
Coding47.2−21.9#94/1652/5
Agentic coding51.8−16.9#60/1571/4
Code generation39.8−30.2#62/771/2
Human preference54.5−13#146/3421/1
Human preference54.5−13#146/3421/1
Reasoning39.8−27.1#154/1782/4
Reasoning46.1−23#76/1401/3
Science38.3−25.2#101/1221/1
Mathematics35−30.7#130/1401/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Claude 3.5 Sonnet, left for the other.
§ 3 · Sources
Where the numbers come from
8 publications, 23 figures. Every one links to the page it was read from.
LMArena Text 1374
LiveCodeBench 36.4%
Aider polyglot 51.6%
MMMU-Pro 51.5%
SimpleBench 41.4%
Vals · CorpFin 53.61%Vals · TaxEval 70.16%Vals · MortgageTax 64.07%Vals · MedQA 83.19%Vals · GPQA Diamond 59.34%Vals · MMLU Pro 78.4%Vals · AIME 10%
HELM Safety · HarmBench 98.1%HELM Safety · SimpleSafetyTests 100%HELM Safety · Anthropic Red Team 99.8%HELM Safety · BBQ 94.9%HELM Safety · XSTest 95.6%
Enkrypt · Jailbreak risk 0.2%Enkrypt · Harmful content risk 0%Enkrypt · CBRN risk 1.8%Enkrypt · Toxicity risk 0.1%Enkrypt · Bias risk 35.4%Enkrypt · Insecure code risk 32.9%
Badge
[](https://publicai.io/model-index/m/claude-3-5-sonnet)