‹ PublicAI Index
The LLM benchmark aggregator.
Claude Sonnet 5.5
Anthropic
Strongest in Biology research (#1 of 23), weakest in Legal (#52 of 157). Above par in 23 of 24 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5.1 and Claude Opus 5.5 and ahead of Kimi K3 and GPT-5.5.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents67.4−0.1#3/2772/5
Knowledge work72.6−0.1#2/1872/2
Reasoning61.4−4.8#9/1833/4
Mathematics64.5−0.4#3/1452/2
Reasoning59.2−9.3#25/1412/3
Coding60.4−8.1#10/1682/5
Code generation70leads#2/781/2
Agentic coding57−11#27/1612/4
Core abilities59.8−7.3#10/2122/3
General intelligence64.7−5.2#9/2122/3
Data analysis58.3−7.3#17/581/1
Language57.4−13.6#17/581/1
Instruction following54.1−19.4#25/581/1
Professional57.7−6.4#12/1771/1
Biology research69.6leads#1/231/1
Medical62.3−4.3#5/1461/1
IT operations57.1−16.6#6/161/1
Finance57.2−5.4#26/1571/1
Cybersecurity46.1−16.4#32/431/1
Legal53.4−13.3#52/1571/1
Knowledge58.2−5#14/1841/3
Document parsing58.7−5.4#12/701/1
Human preference63.6−5.2#40/3451/1
Human preference63.6−5.2#40/3451/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Claude Sonnet 5.5, left for the other.
§ 3 · Sources
Where the numbers come from
8 publications, 30 figures. Every one links to the page it was read from.
LMArena Text 1471
Artificial Analysis Intelligence Index 56
GDPval-AA 1840
AA-Briefcase 1824
LiveBench 77.8LiveBench · Reasoning 86.8LiveBench · Coding 88.9LiveBench · Agentic Coding 39.3LiveBench · Mathematics 96.7LiveBench · Data Analysis 78.6LiveBench · Language 83.4LiveBench · Instruction Following 70.5
SimpleBench 75.9%
Vals · Legal Research Bench 48.08%Vals · Harvey Legal Agent Benchmark 2.92%Vals · Finance Agent 58.1%Vals · MedCode 52.92%Vals · MedScribe 91.1%Vals · BioMysteryBench 81.11%Vals · CyberBench 59.58%Vals · SRE Bench 30.15%Vals · Vibe Code Bench 92.39% (OpenHands)Vals · Code Migration 69.83%Vals · ProofBench 100%
ParseBench · Overall 70.18%ParseBench · Tables 91.04%ParseBench · Charts 37.69%ParseBench · Content faithfulness 91.64%ParseBench · Semantic formatting 71.63%ParseBench · Visual grounding 58.92%
Badge
[](https://publicai.io/model-index/m/claude-sonnet-5-5)