‹ PublicAI Index
The LLM benchmark aggregator.
Claude Opus 5.5
Anthropic
Strongest in Human preference (#1 of 342), weakest in Toxicity avoidance (#119 of 272). Above par in 26 of 27 scopes. Among the models it meets almost everywhere, it finishes behind none and ahead of Claude Fable 5 and Claude Fable 5.1.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Human preference67.5leads#1/3421/1
Human preference67.5leads#1/3421/1
Agents68.1leads#1/2682/5
Knowledge work73.6leads#1/1782/2
Reasoning66.9leads#1/1784/4
Reasoning69.1leads#1/1403/3
Mathematics65.7leads#1/1402/2
Coding69.1leads#1/1652/5
Agentic coding68.7leads#1/1572/4
Code generation70leads#1/771/2
Safety60.6−0.9#2/3371/3
Secure code66.3−0.2#2/2741/1
Fairness70leads#3/3001/2
Harm refusal58.2−3.4#23/3001/2
Jailbreak resistance59.8−5.1#49/2721/1
Toxicity avoidance54.6−4.4#119/2721/1
Core abilities62.2−5.2#6/2042/3
General intelligence70.1leads#1/2042/3
Data analysis61.6−4.5#6/571/1
Language63.2−8.3#8/571/1
Instruction following46.1−27.6#35/571/1
Professional59.7−1.9#7/1681/1
Biology research68.5leads#1/161/1
IT operations62.3−11.5#2/111/1
Medical61.1−5.7#5/1401/1
Finance58.1−4.8#18/1511/1
Legal55.5−11.6#38/1511/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Claude Opus 5.5, left for the other.
§ 3 · Sources
Where the numbers come from
9 publications, 30 figures. Every one links to the page it was read from.
LMArena Text 1509
Artificial Analysis Intelligence Index 58
GDPval-AA 1846
AA-Briefcase 1822
ARC-AGI-2 91.7%
LiveBench 83.2LiveBench · Reasoning 92.2LiveBench · Coding 89.3LiveBench · Agentic Coding 71.7LiveBench · Mathematics 97.1LiveBench · Data Analysis 80.3LiveBench · Language 86.3LiveBench · Instruction Following 65.7
SimpleBench 88.4%
Vals · Legal Research Bench 50.48%Vals · Harvey Legal Agent Benchmark 3.75%Vals · Finance Agent 58.59%Vals · MedCode 49.8%Vals · MedScribe 91.43%Vals · BioMysteryBench 79.26%Vals · SRE Bench 33.59%Vals · Vibe Code Bench 90.29% (OpenHands)Vals · Code Migration 66.65%Vals · ProofBench 100%
Enkrypt · Jailbreak risk 4.5%Enkrypt · Harmful content risk 1.7%Enkrypt · CBRN risk 10.2%Enkrypt · Toxicity risk 3.1%Enkrypt · Bias risk 25.6%Enkrypt · Insecure code risk 0.4%
Badge
[](https://publicai.io/model-index/m/claude-opus-5-5)