‹ PublicAI Index
The LLM benchmark aggregator.
Claude 3 Sonnet
Anthropic
Strongest in Harm refusal (#6 of 300), weakest in Human preference (#236 of 342). Above par in 6 of 9 scopes. Among the models it meets almost everywhere, it finishes behind Claude 3.5 Sonnet and GPT-5 Nano and ahead of Gemini 2.5 Pro and DeepSeek R1.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety55.5−6#42/3372/3
Harm refusal60.2−1.4#6/3002/2
Jailbreak resistance62.8−2.1#15/2721/1
Toxicity avoidance58.4−0.6#18/2721/1
Safe-prompt compliance31.5−28.9#76/821/1
Secure code56.7−9.8#103/2741/1
Fairness50.4−19.6#109/3002/2
Human preference45.5−22#236/3421/1
Human preference45.5−22#236/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Claude 3 Sonnet, left for the other.
§ 3 · Sources
Where the numbers come from
3 publications, 12 figures. Every one links to the page it was read from.
LMArena Text 1281
HELM Safety · HarmBench 95.8%HELM Safety · SimpleSafetyTests 100%HELM Safety · Anthropic Red Team 99.8%HELM Safety · BBQ 90%HELM Safety · XSTest 85.8%
Enkrypt · Jailbreak risk 2%Enkrypt · Harmful content risk 6.7%Enkrypt · CBRN risk 3.7%Enkrypt · Toxicity risk 0.4%Enkrypt · Bias risk 78.3%Enkrypt · Insecure code risk 20%
Badge
[](https://publicai.io/model-index/m/claude-3-sonnet)