‹ PublicAI Index
The LLM benchmark aggregator.
Claude Opus 4
Anthropic
Strongest in Jailbreak resistance (#14 of 272), weakest in Secure code (#136 of 274). Above par in 20 of 25 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Gemini 2.5 Pro and GPT-5.6 Luna.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Core abilities57.3−10.1#24/2041/3
General intelligence60.6−9.5#25/2041/3
Coding54.5−14.6#35/1652/5
Agentic coding57.1−11.6#27/1571/4
Code generation50.4−19.6#43/771/2
Safety56.4−5.1#36/3373/3
Jailbreak resistance62.8−2.1#14/2721/1
Safe-prompt compliance56.4−4#25/821/1
Fairness62.5−7.5#28/3002/2
Toxicity avoidance57.8−1.2#35/2721/1
Harm refusal55.6−6#63/3002/2
Factual grounding44.9−25.7#77/1011/1
Secure code51.8−14.7#136/2741/1
Knowledge55.2−8.3#51/1381/2
Academic knowledge56.2−7.7#47/1231/1
Professional52.4−9.2#56/1681/1
Medical55−11.8#37/1401/1
Legal53.7−13.4#48/1511/1
Finance50.6−12.3#84/1511/1
Human preference59.5−8#94/3421/1
Human preference59.5−8#94/3421/1
Reasoning47−19.9#110/1783/4
Reasoning49.8−19.3#62/1402/3
Science47−16.5#86/1221/1
Mathematics43.4−22.3#104/1401/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Claude Opus 4, left for the other.
§ 3 · Sources
Where the numbers come from
10 publications, 25 figures. Every one links to the page it was read from.
LMArena Text 1426
ARC-AGI-2 8.6%
LiveCodeBench 56.6%
Aider polyglot 72%
Kagi LLM Benchmark 74.3%
SimpleBench 58.8%
Vals · LegalBench 83.07%Vals · TaxEval 71.91%Vals · MortgageTax 58.59%Vals · MedQA 92.87%Vals · GPQA Diamond 71.72%Vals · MMLU Pro 86.17%Vals · AIME 41.25%
HELM Safety · HarmBench 91.7%HELM Safety · SimpleSafetyTests 100%HELM Safety · Anthropic Red Team 98.9%HELM Safety · BBQ 97.1%HELM Safety · XSTest 97%
Enkrypt · Jailbreak risk 2%Enkrypt · Harmful content risk 1.1%Enkrypt · CBRN risk 25%Enkrypt · Toxicity risk 0.8%Enkrypt · Bias risk 56.6%Enkrypt · Insecure code risk 29.8%
Vectara · Factual consistency 88%
Badge
[](https://publicai.io/model-index/m/claude-opus-4)