‹ PublicAI Index
The LLM benchmark aggregator.
Claude Opus 4.8
Anthropic
Strongest in Medical (#6 of 140), weakest in Secure code (#121 of 274). Above par in 27 of 28 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Grok 4.7 and GPT-5.5.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety59.8−1.7#8/3371/3
Fairness69.1−0.9#14/3001/2
Harm refusal59.2−2.4#16/3001/2
Jailbreak resistance61.6−3.3#27/2721/1
Toxicity avoidance58−1#32/2721/1
Secure code54.3−12.2#121/2741/1
Professional58.6−3#10/1681/1
Medical60.5−6.3#6/1401/1
Finance59.1−3.8#8/1511/1
Legal58.4−8.7#21/1511/1
Knowledge58.7−4.8#11/1381/2
Academic knowledge60.4−3.5#9/1231/1
Reasoning59.6−7.3#15/1784/4
Science61.4−2.1#16/1221/1
Reasoning60.1−9#21/1403/3
Mathematics57.5−8.2#22/1401/2
Human preference64.7−2.8#19/3421/1
Human preference64.7−2.8#19/3421/1
Coding57.5−11.6#20/1653/5
Code generation59−11#14/771/2
Agentic coding57.1−11.6#25/1573/4
Agents60.3−7.8#30/2682/5
Knowledge work63.4−10.2#26/1782/2
Core abilities54.4−13#53/2042/3
General intelligence63.3−6.8#9/2042/3
Instruction following57.2−16.5#17/571/1
Language50.8−20.7#30/571/1
Data analysis37.5−28.6#49/571/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Claude Opus 4.8, left for the other.
§ 3 · Sources
Where the numbers come from
10 publications, 35 figures. Every one links to the page it was read from.
LMArena Text 1480
GDPval-AA 1438
AA-Briefcase 1321
Terminal-Bench 23.6% (Claude Code)
ARC-AGI-2 72.1%
LiveBench 76.2LiveBench · Reasoning 89.2LiveBench · Coding 81.8LiveBench · Agentic Coding 50.5LiveBench · Mathematics 94.3LiveBench · Data Analysis 66LiveBench · Language 79.7LiveBench · Instruction Following 72
Kagi LLM Benchmark 88.8%
SimpleBench 64.8%
Vals · Legal Research Bench 43.75%Vals · LegalBench 83.57%Vals · Harvey Legal Agent Benchmark 9.58%Vals · Finance Agent 53.92%Vals · CorpFin 66.71%Vals · TaxEval 75.63%Vals · MortgageTax 69.91%Vals · MedCode 53.22%Vals · MedScribe 85.75%Vals · SWE-bench Verified 88.6% (Mini-SWE-agent)Vals · Vibe Code Bench 82.72% (OpenHands)Vals · Code Migration 47.25%Vals · GPQA Diamond 92.42%Vals · MMLU Pro 89.58%
Enkrypt · Jailbreak risk 3%Enkrypt · Harmful content risk 0.6%Enkrypt · CBRN risk 8%Enkrypt · Toxicity risk 0.7%Enkrypt · Bias risk 48.1%Enkrypt · Insecure code risk 24.9%
Badge
[](https://publicai.io/model-index/m/claude-opus-4-8)