‹ PublicAI Index
The LLM benchmark aggregator.
Claude Sonnet 4.6
Anthropic
Strongest in Computer use (#8 of 37), weakest in Secure code (#154 of 274). Above par in 22 of 28 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5.5 and Claude Fable 5 and ahead of O3 and GPT-5.1.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Professional55.8−5.8#30/1681/1
Finance57.7−5.2#22/1511/1
Legal55.1−12#41/1511/1
Medical54.6−12.2#45/1401/1
Knowledge56.4−7.1#32/1381/2
Academic knowledge57.6−6.3#29/1231/1
Human preference63.9−3.6#35/3421/1
Human preference63.9−3.6#35/3421/1
Agents58.8−9.3#40/2683/5
Computer use64.2−7.5#8/371/1
Knowledge work57.9−15.7#41/1782/2
Reasoning53.7−13.2#51/1783/4
Reasoning53.7−15.4#46/1402/3
Science56.6−6.9#48/1221/1
Mathematics52.4−13.3#70/1402/2
Coding51.3−17.8#65/1652/5
Code generation53.9−16.1#31/771/2
Agentic coding50.6−18.1#69/1572/4
Core abilities47.6−19.8#120/2041/3
Data analysis57.6−8.5#22/571/1
Language44−27.5#39/571/1
Instruction following41.7−32#43/571/1
General intelligence47.1−23#123/2041/3
Safety52−9.5#131/3372/3
Harm refusal56.5−5.1#42/3001/2
Factual grounding48.4−22.2#63/1011/1
Fairness47.9−22.1#145/3001/2
Secure code50.3−16.2#154/2741/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Claude Sonnet 4.6, left for the other.
§ 3 · Sources
Where the numbers come from
9 publications, 33 figures. Every one links to the page it was read from.
LMArena Text 1472
GDPval-AA 1220
AA-Briefcase 1063
ARC-AGI-2 58.3%
LiveBench 73LiveBench · Reasoning 84.8LiveBench · Coding 79.3LiveBench · Agentic Coding 42.6LiveBench · Mathematics 87LiveBench · Data Analysis 77.9LiveBench · Language 76.1LiveBench · Instruction Following 63.2
OSWorld 72.1%
Vals · Legal Research Bench 38.46%Vals · CaseLaw 63.99%Vals · LegalBench 82.12%Vals · Harvey Legal Agent Benchmark 5%Vals · Finance Agent 51.03%Vals · CorpFin 65.31%Vals · TaxEval 77.11%Vals · MortgageTax 67.73%Vals · MedQA 92.06%Vals · SWE-bench Verified 77.4% (Mini-SWE-agent)Vals · Vibe Code Bench 55.77% (Claude Code)Vals · Code Migration 39.89%Vals · GPQA Diamond 85.61%Vals · MMLU Pro 87.34%Vals · AIME 92.29%
Enkrypt · Harmful content risk 0.6%Enkrypt · CBRN risk 15.3%Enkrypt · Bias risk 81.1%Enkrypt · Insecure code risk 32.9%
Vectara · Factual consistency 89.4%
Badge
[](https://publicai.io/model-index/m/claude-sonnet-4-6)