‹ PublicAI Index
The LLM benchmark aggregator.
Claude Opus 4.6
Anthropic
Strongest in Human preference (#2 of 342), weakest in Secure code (#196 of 274). Above par in 21 of 25 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Mimo V2.6 Pro and Claude Sonnet 5.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Human preference67.1−0.4#2/3421/1
Human preference67.1−0.4#2/3421/1
Professional57.4−4.2#14/1681/1
Finance58.7−4.2#10/1511/1
Medical58.9−7.9#13/1401/1
Legal55.8−11.3#36/1511/1
Knowledge58.2−5.3#17/1381/2
Academic knowledge59.9−4#15/1231/1
Reasoning58.1−8.8#21/1784/4
Reasoning60.1−9#20/1403/3
Science59.5−4#28/1221/1
Mathematics55.1−10.6#48/1402/2
Core abilities52.7−14.7#65/2042/3
Language57.6−13.9#17/571/1
General intelligence60−10.1#27/2042/3
Instruction following41.9−31.8#42/571/1
Data analysis44.1−22#45/571/1
Coding51.1−18#66/1652/5
Code generation51.6−18.4#39/771/2
Agentic coding50.9−17.8#66/1572/4
Safety52−9.5#133/3372/3
Harm refusal59−2.6#18/3001/2
Factual grounding44.4−26.2#80/1011/1
Fairness50.3−19.7#110/3001/2
Secure code44.7−21.8#196/2741/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Claude Opus 4.6, left for the other.
§ 3 · Sources
Where the numbers come from
8 publications, 30 figures. Every one links to the page it was read from.
LMArena Text 1505
ARC-AGI-2 68.8%
LiveBench 74.5LiveBench · Reasoning 88.7LiveBench · Coding 78.2LiveBench · Agentic Coding 49LiveBench · Mathematics 89.3LiveBench · Data Analysis 69.9LiveBench · Language 83.3LiveBench · Instruction Following 63.3
Kagi LLM Benchmark 83.6%
SimpleBench 67.6%
Vals · CaseLaw 62.06%Vals · LegalBench 85.3%Vals · CorpFin 67.02%Vals · TaxEval 75.96%Vals · MortgageTax 68.52%Vals · MedQA 95.41%Vals · MedCode 49.13%Vals · MedScribe 86.13%Vals · SWE-bench Verified 78.2% (Mini-SWE-agent)Vals · Vibe Code Bench 53.5% (OpenHands)Vals · GPQA Diamond 89.65%Vals · MMLU Pro 89.11%Vals · AIME 95.63%
Enkrypt · Harmful content risk 0.6%Enkrypt · CBRN risk 8.7%Enkrypt · Bias risk 77.3%Enkrypt · Insecure code risk 44.4%
Vectara · Factual consistency 87.8%
Badge
[](https://publicai.io/model-index/m/claude-opus-4-6)