‹ PublicAI Index
The LLM benchmark aggregator.
Claude 3 Opus
Anthropic
Strongest in Harm refusal (#7 of 300), weakest in Human preference (#201 of 342). Above par in 6 of 11 scopes. Among the models it meets almost everywhere, it finishes behind Claude 3.5 Sonnet and Claude Opus 5.5 and ahead of Gemini 2.5 Pro and Claude Opus 4.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety59.6−1.9#10/3372/3
Harm refusal60.1−1.5#7/3002/2
Jailbreak resistance63.4−1.5#7/2721/1
Toxicity avoidance58.8−0.2#9/2721/1
Fairness64.7−5.3#25/3002/2
Safe-prompt compliance46.4−14#67/821/1
Secure code59.1−7.4#74/2741/1
Reasoning42.5−24.4#142/1781/4
Reasoning38−31.1#122/1401/3
Human preference49.5−18#201/3421/1
Human preference49.5−18#201/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Claude 3 Opus, left for the other.
§ 3 · Sources
Where the numbers come from
4 publications, 13 figures. Every one links to the page it was read from.
LMArena Text 1322
SimpleBench 23.5%
HELM Safety · HarmBench 97.4%HELM Safety · SimpleSafetyTests 100%HELM Safety · Anthropic Red Team 99.8%HELM Safety · BBQ 94%HELM Safety · XSTest 92.5%
Enkrypt · Jailbreak risk 1.5%Enkrypt · Harmful content risk 7.8%Enkrypt · CBRN risk 4.3%Enkrypt · Toxicity risk 0.1%Enkrypt · Bias risk 25.6%Enkrypt · Insecure code risk 15.1%
Badge
[](https://publicai.io/model-index/m/claude-3-opus)