‹ PublicAI Index
The LLM benchmark aggregator.
Claude Opus 4.1
Anthropic
Strongest in Academic knowledge (#23 of 123), weakest in Safety (#223 of 337). Above par in 12 of 14 scopes. Among the models it meets almost everywhere, it finishes behind GPT-5.5 and Claude Opus 5 and ahead of GLM-5.2 and GPT-5.6 Luna.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge57−6.5#26/1381/2
Academic knowledge58.4−5.5#23/1231/1
Professional52.3−9.3#57/1681/1
Legal54−13.1#46/1511/1
Medical53.4−13.4#57/1401/1
Finance50.5−12.4#87/1511/1
Human preference61.8−5.7#61/3421/1
Human preference61.8−5.7#61/3421/1
Reasoning52.8−14.1#67/1782/4
Reasoning54.4−14.7#44/1401/3
Mathematics53.2−12.5#64/1401/2
Science50.1−13.4#76/1221/1
Safety48.6−12.9#223/3371/3
Factual grounding45.4−25.2#73/1011/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Claude Opus 4.1, left for the other.
§ 3 · Sources
Where the numbers come from
4 publications, 12 figures. Every one links to the page it was read from.
LMArena Text 1450
SimpleBench 60%
Vals · LegalBench 83.46%Vals · TaxEval 73.67%Vals · MortgageTax 56.12%Vals · MedQA 93.59%Vals · MedCode 47.23%Vals · MedScribe 73.9%Vals · GPQA Diamond 76.26%Vals · MMLU Pro 87.92%Vals · AIME 78.18%
Vectara · Factual consistency 88.2%
Badge
[](https://publicai.io/model-index/m/claude-opus-4-1)