‹ PublicAI Index
The LLM benchmark aggregator.
Claude Opus 4.5
Anthropic
Strongest in Tool use (#1 of 81), weakest in Safety (#203 of 337). Above par in 19 of 24 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5 and Claude Fable 5 and ahead of Claude Sonnet 4.6 and Grok 4.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents63.3−4.8#13/2681/5
Tool use74leads#1/811/1
Professional56.9−4.7#19/1681/1
Medical58.8−8#15/1401/1
Finance57.2−5.7#27/1511/1
Legal55.8−11.3#35/1511/1
Human preference64.1−3.4#31/3421/1
Human preference64.1−3.4#31/3421/1
Knowledge56.3−7.2#34/1381/2
Academic knowledge57.5−6.4#31/1231/1
Reasoning53.6−13.3#54/1784/4
Mathematics56.1−9.6#37/1402/2
Science56.8−6.7#45/1221/1
Reasoning50.3−18.8#60/1403/3
Core abilities52−15.4#71/2042/3
Language53.8−17.7#23/571/1
Data analysis51.7−14.4#31/571/1
Instruction following40.5−33.2#45/571/1
General intelligence57−13.1#46/2042/3
Coding46.4−22.7#106/1652/5
Code generation54.7−15.3#29/771/2
Agentic coding43.6−25.1#117/1572/4
Safety49.3−12.2#203/3371/3
Factual grounding47.7−22.9#68/1011/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Claude Opus 4.5, left for the other.
§ 3 · Sources
Where the numbers come from
8 publications, 27 figures. Every one links to the page it was read from.
LMArena Text 1474
ARC-AGI-2 37.6%
LiveBench 72.6LiveBench · Reasoning 80.1LiveBench · Coding 79.7LiveBench · Agentic Coding 39.7LiveBench · Mathematics 90.4LiveBench · Data Analysis 74.4LiveBench · Language 81.3LiveBench · Instruction Following 62.5
BFCL v4 77.47%
Kagi LLM Benchmark 80.2%
SimpleBench 62%
Vals · CaseLaw 62.59%Vals · LegalBench 84.6%Vals · CorpFin 65.07%Vals · TaxEval 74.86%Vals · MortgageTax 67.69%Vals · MedQA 95.88%Vals · MedCode 49.16%Vals · MedScribe 85.32%Vals · SWE-bench Verified 76.4% (Mini-SWE-agent)Vals · Vibe Code Bench 20.63% (OpenHands)Vals · GPQA Diamond 85.86%Vals · MMLU Pro 87.26%Vals · AIME 95.42%
Vectara · Factual consistency 89.1%
Badge
[](https://publicai.io/model-index/m/claude-opus-4-5)