Request a pilot
‹ PublicAI Index

The LLM benchmark aggregator.

Claude Opus 5

Anthropic

Strongest in Finance (#1 of 151), weakest in Jailbreak resistance (#166 of 272). Above par in 34 of 36 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5.5 and Claude Fable 5.1 and ahead of GPT-6 Astra and Grok 4.7.

§ 1 · Profile

What it is good at

Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.

Professional61.3#2/168
Finance62.9#1/151
Medical66.8#1/140
Biology research68#3/16
IT operations47.9#5/11
Cybersecurity50.7#7/9
Legal60.6#9/151
Agents68#2/268
Computer use70.3#2/37
Knowledge work70.5#3/178
Workflow automation55.3—/0✱
Knowledge60.8#3/138
Academic knowledge63#2/123
Multimodal understanding60.1—/19✱
Coding65.2#4/165
Agentic coding67#2/157
Code generation58.2#16/77
Repository Q&A66.7—/0✱
Reasoning64.9#4/178
Reasoning66.5#4/140
Mathematics64.3#4/140
Science62.1#7/122
Expert reasoning55.9—/0✱
Human preference65.5#11/342
Human preference65.5#11/342
Safety57.6#21/337
Secure code66.3#3/274
Fairness70#9/300
Toxicity avoidance55.8#86/272
Harm refusal52.7#116/300
Jailbreak resistance51.5#166/272
Core abilities55.4#41/204
Language67.7#4/57
Data analysis52#29/57
General intelligence58.5#37/204
Instruction following42.8#41/57

§ 2 · Head to head

What it beats, and what beats it

The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Claude Opus 5, left for the other.

Qwen3.6 Pluswon 22 · lost 0
GPT-5.4 Miniwon 22 · lost 0
Gemma 4 31Bwon 22 · lost 0
GLM-5.1won 22 · lost 0
Claude Opus 4.5won 22 · lost 0
GLM-4.6won 23 · lost 0
Gemini 2.5 Flashwon 24 · lost 0
DeepSeek V3.2won 24 · lost 0
DeepSeek V3won 25 · lost 0
GPT-6 Lunawon 26 · lost 0
Grok 4.3won 28 · lost 0
GLM-5.2won 31 · lost 1
Kimi K2.6won 28 · lost 1
Gemini 3.6 Flashwon 28 · lost 1
GLM-5.3 Flashwon 27 · lost 1
Qwen3.7 Maxwon 27 · lost 1
Gemini 3.5 Flashwon 27 · lost 1
GPT OSS 120Bwon 23 · lost 1
GPT-5.4 Nanowon 22 · lost 1
Kimi K2won 22 · lost 1
Mimo V2.5 Prowon 22 · lost 1
Command Awon 21 · lost 1
Gemini 2.0 Flashwon 21 · lost 1
Gemini 3.5 Flash Litewon 21 · lost 1
DeepSeek V4 Flashwon 31 · lost 2
MiniMax M3won 30 · lost 2
Claude Sonnet 4.6won 25 · lost 2
Kimi K2.5won 23 · lost 2
GPT-4.1 Nanowon 22 · lost 2
GPT-4.1 Miniwon 22 · lost 2
O3 Miniwon 21 · lost 2
Grok 4won 21 · lost 2
GPT-4.1won 21 · lost 2
Qwen3 235B A22B Instructwon 21 · lost 2
Qwen3.8 27Bwon 21 · lost 2
Llama 4 Scout Instructwon 20 · lost 2
DeepSeek V4 Prowon 29 · lost 3
GLM-5.3won 23 · lost 3
GPT-4o Miniwon 22 · lost 3
Gemini 2.5 Prowon 22 · lost 3
Claude Sonnet 4.5won 22 · lost 3
O4 Miniwon 21 · lost 3
Claude Haiku 4.5won 21 · lost 3
DeepSeek R1won 21 · lost 3
Gemini 3.7 Flashwon 20 · lost 3
Claude 3.7 Sonnetwon 19 · lost 3
GPT-5.1won 19 · lost 3
Kimi K3won 24 · lost 4
O3won 21 · lost 4
Inklingwon 26 · lost 5
GPT-5.6 Lunawon 26 · lost 5
Claude Opus 4.6won 20 · lost 4
GPT-5.6 Terrawon 25 · lost 5
Gemini 3.8 Flashwon 25 · lost 5
Grok 4.5won 24 · lost 5
Gemini 3.1 Prowon 19 · lost 4
GPT-5.4won 23 · lost 5
GPT-4owon 18 · lost 4
GPT-5.2won 22 · lost 5
Claude Sonnet 4won 21 · lost 5
GPT-6 Solwon 21 · lost 5
GPT-5 Miniwon 19 · lost 5
GPT-5won 19 · lost 5
Qwen3.8 Maxwon 22 · lost 6
Claude Opus 4won 18 · lost 5
Muse Spark 1.2won 17 · lost 5
GPT-5.5won 22 · lost 7
Muse Spark 1.1won 22 · lost 7
DeepSeek V4.1 Flashwon 21 · lost 7
Claude Opus 4.8won 21 · lost 7
Claude Sonnet 5won 22 · lost 8
GPT-5.6 Solwon 25 · lost 10
Claude Opus 4.7won 19 · lost 9
Grok 4.6won 20 · lost 10
Grok 4.7won 16 · lost 10
GPT-6 Astrawon 13 · lost 10
Claude Fable 5won 11 · lost 12
Claude Fable 5.1won 6 · lost 18
Claude Opus 5.5won 6 · lost 21

§ 3 · Sources

Where the numbers come from

11 publications, 54 figures. Every one links to the page it was read from.

LMArena Text ↗1 measureread 2026-09-26
LMArena Text 1488
GDPval-AA ↗1 measureread 2026-09-26
GDPval-AA 1708
AA-Briefcase ↗1 measureread 2026-09-26
AA-Briefcase 1673
Terminal-Bench ↗1 measureread 2026-09-26
Terminal-Bench 53.9% (Claude Code)
ARC-AGI-2 ↗1 measureread 2026-09-26
ARC-AGI-2 90.4%
LiveBench ↗8 measuresread 2026-09-26
LiveBench 80.1LiveBench · Reasoning 91.2LiveBench · Coding 81.4LiveBench · Agentic Coding 65.2LiveBench · Mathematics 95.7LiveBench · Data Analysis 74.6LiveBench · Language 88.7LiveBench · Instruction Following 63.8
OSWorld ↗1 measureread 2026-09-26
OSWorld 83.4%
SimpleBench ↗1 measureread 2026-09-26
SimpleBench 80.6%
Vals.ai ↗18 measuresread 2026-09-26
Vals · Legal Research Bench 55.29%Vals · LegalBench 86.97%Vals · Harvey Legal Agent Benchmark 6.67%Vals · Finance Agent 58.63%Vals · CorpFin 73.19%Vals · TaxEval 75.14%Vals · MortgageTax 72.06%Vals · MedCode 63.57%Vals · MedScribe 90.98%Vals · BioMysteryBench 79.26%Vals · CyberBench 65.36%Vals · SRE Bench 12.21%Vals · SWE-bench Verified 97% (Mini-SWE-agent)Vals · Vibe Code Bench 88.4% (OpenHands)Vals · Code Migration 57.47%Vals · GPQA Diamond 93.43%Vals · MMLU Pro 91.59%Vals · ProofBench 99%
Enkrypt AI Safety Leaderboard ↗6 measuresread 2026-09-26
Enkrypt · Jailbreak risk 11.5%Enkrypt · Harmful content risk 0.6%Enkrypt · CBRN risk 25.2%Enkrypt · Toxicity risk 2.2%Enkrypt · Bias risk 40.8%Enkrypt · Insecure code risk 0.4%
GPQA Diamond (Pass@1) 93.4%HLE (Pass@1) 56.3%Terminal-Bench 2.1 (Pass@1) 89.1%Terminal-Bench 3.0 (Pass@1) 43.3%Terminal-Bench 4.0 (Pass@1) 51.8%DeepSWE v1.1 (Resolved) 74%ProgramBench (Almost@1) 37%NL2Repo-Bench (Score) 75.3%ExploitGym (Pass@1) 22.1%HLE w/ tools (Pass@1) 63.6%AutomationBench (Pass@1) 50.3%Agent's Last Exam (Pass@1) 28.6%Chartography w/ tools (Pass@1) 84%BabyVision w/ tools (Pass@1) 94.1%ZeroBench-main w/ tools (Pass@5) 52%

Badge

PublicAI Index badge for Claude Opus 5[![PublicAI Index](https://publicai.io/model-index/badge?model=claude-opus-5)](https://publicai.io/model-index/m/claude-opus-5)

✱ Placed by a one-off publication, not a board that re-ran the model. ✱✱ marks a figure the model’s own publisher printed, which counts for less again. Scores are 0–100 on the PublicAI Index scale, 50 = the average of the models each source lists. Scores belong to their publishers.

Snapshot 2026-09-26 · the full index · JSON API