‹ PublicAI Index
The LLM benchmark aggregator.
Claude Opus 5
Anthropic
Strongest in Finance (#1 of 151), weakest in Jailbreak resistance (#166 of 272). Above par in 34 of 36 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5.5 and Claude Fable 5.1 and ahead of GPT-6 Astra and Grok 4.7.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Professional61.3−0.3#2/1681/1
Finance62.9leads#1/1511/1
Medical66.8leads#1/1401/1
Biology research68−0.5#3/161/1
IT operations47.9−25.9#5/111/1
Cybersecurity50.7−9#7/91/1
Legal60.6−6.5#9/1511/1
Agents68−0.1#2/2683/5
Computer use70.3−1.4#2/371/1
Knowledge work70.5−3.1#3/1782/2
Workflow automation55.3−6.7—/0✱✱
Knowledge60.8−2.7#3/1381/2
Academic knowledge63−0.9#2/1231/1
Multimodal understanding60.1−12—/19✱0/1
Coding65.2−3.9#4/1653/5
Agentic coding67−1.7#2/1573/4
Code generation58.2−11.8#16/771/2
Repository Q&A66.7leads—/0✱✱
Reasoning64.9−2#4/1784/4
Reasoning66.5−2.6#4/1403/3
Mathematics64.3−1.4#4/1402/2
Science62.1−1.4#7/1221/1
Expert reasoning55.9−5—/0✱✱
Human preference65.5−2#11/3421/1
Human preference65.5−2#11/3421/1
Safety57.6−3.9#21/3371/3
Secure code66.3−0.2#3/2741/1
Fairness70leads#9/3001/2
Toxicity avoidance55.8−3.2#86/2721/1
Harm refusal52.7−8.9#116/3001/2
Jailbreak resistance51.5−13.4#166/2721/1
Core abilities55.4−12#41/2041/3
Language67.7−3.8#4/571/1
Data analysis52−14.1#29/571/1
General intelligence58.5−11.6#37/2041/3
Instruction following42.8−30.9#41/571/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Claude Opus 5, left for the other.
§ 3 · Sources
Where the numbers come from
11 publications, 54 figures. Every one links to the page it was read from.
LMArena Text 1488
GDPval-AA 1708
AA-Briefcase 1673
Terminal-Bench 53.9% (Claude Code)
ARC-AGI-2 90.4%
LiveBench 80.1LiveBench · Reasoning 91.2LiveBench · Coding 81.4LiveBench · Agentic Coding 65.2LiveBench · Mathematics 95.7LiveBench · Data Analysis 74.6LiveBench · Language 88.7LiveBench · Instruction Following 63.8
OSWorld 83.4%
SimpleBench 80.6%
Vals · Legal Research Bench 55.29%Vals · LegalBench 86.97%Vals · Harvey Legal Agent Benchmark 6.67%Vals · Finance Agent 58.63%Vals · CorpFin 73.19%Vals · TaxEval 75.14%Vals · MortgageTax 72.06%Vals · MedCode 63.57%Vals · MedScribe 90.98%Vals · BioMysteryBench 79.26%Vals · CyberBench 65.36%Vals · SRE Bench 12.21%Vals · SWE-bench Verified 97% (Mini-SWE-agent)Vals · Vibe Code Bench 88.4% (OpenHands)Vals · Code Migration 57.47%Vals · GPQA Diamond 93.43%Vals · MMLU Pro 91.59%Vals · ProofBench 99%
Enkrypt · Jailbreak risk 11.5%Enkrypt · Harmful content risk 0.6%Enkrypt · CBRN risk 25.2%Enkrypt · Toxicity risk 2.2%Enkrypt · Bias risk 40.8%Enkrypt · Insecure code risk 0.4%
GPQA Diamond (Pass@1) 93.4%HLE (Pass@1) 56.3%Terminal-Bench 2.1 (Pass@1) 89.1%Terminal-Bench 3.0 (Pass@1) 43.3%Terminal-Bench 4.0 (Pass@1) 51.8%DeepSWE v1.1 (Resolved) 74%ProgramBench (Almost@1) 37%NL2Repo-Bench (Score) 75.3%ExploitGym (Pass@1) 22.1%HLE w/ tools (Pass@1) 63.6%AutomationBench (Pass@1) 50.3%Agent's Last Exam (Pass@1) 28.6%Chartography w/ tools (Pass@1) 84%BabyVision w/ tools (Pass@1) 94.1%ZeroBench-main w/ tools (Pass@5) 52%
Badge
[](https://publicai.io/model-index/m/claude-opus-5)