Request a pilot
‹ PublicAI Index

The LLM benchmark aggregator.

Claude Sonnet 5.5

Anthropic

Strongest in Biology research (#1 of 23), weakest in Legal (#52 of 157). Above par in 23 of 24 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5.1 and Claude Opus 5.5 and ahead of Kimi K3 and GPT-5.5.

§ 1 · Profile

What it is good at

Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.

Agents67.4#3/277
Knowledge work72.6#2/187
Reasoning61.4#9/183
Mathematics64.5#3/145
Reasoning59.2#25/141
Coding60.4#10/168
Code generation70#2/78
Agentic coding57#27/161
Core abilities59.8#10/212
General intelligence64.7#9/212
Data analysis58.3#17/58
Language57.4#17/58
Instruction following54.1#25/58
Professional57.7#12/177
Biology research69.6#1/23
Medical62.3#5/146
IT operations57.1#6/16
Finance57.2#26/157
Cybersecurity46.1#32/43
Legal53.4#52/157
Knowledge58.2#14/184
Document parsing58.7#12/70
Human preference63.6#40/345
Human preference63.6#40/345

§ 2 · Head to head

What it beats, and what beats it

The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Claude Sonnet 5.5, left for the other.

GPT-4owon 15 · lost 0
O3 Miniwon 15 · lost 0
Kimi K2won 15 · lost 0
Qwen3 235B A22B Instructwon 15 · lost 0
GLM-4.6won 15 · lost 0
Mistral Medium 3.5won 15 · lost 0
Gemma 4 31Bwon 15 · lost 0
Qwen3.6 27Bwon 16 · lost 0
GPT-4o Miniwon 16 · lost 0
GPT-4.1 Nanowon 16 · lost 0
GPT OSS 120Bwon 16 · lost 0
GPT-4.1 Miniwon 16 · lost 0
Gemini 2.5 Flashwon 16 · lost 0
Mistral Large 3won 16 · lost 0
DeepSeek V3.2won 16 · lost 0
Gemini 3.1 Flash Litewon 16 · lost 0
Kimi K2.5won 16 · lost 0
DeepSeek V3won 17 · lost 0
Claude Sonnet 4won 17 · lost 0
Claude Haiku 4.5won 17 · lost 0
DeepSeek R1won 17 · lost 0
GPT-5.4 Miniwon 19 · lost 0
Grok 4.3won 20 · lost 0
Qwen3.6 Pluswon 20 · lost 0
Kimi K2.6won 20 · lost 0
GPT-5.4 Nanowon 21 · lost 0
Gemini 3.5 Flash Litewon 22 · lost 0
GPT-6 Lunawon 22 · lost 1
MiniMax M3won 20 · lost 1
Inklingwon 20 · lost 1
GPT-5 Miniwon 16 · lost 1
Inkling Smallwon 16 · lost 1
Gemini 2.5 Prowon 16 · lost 1
O4 Miniwon 15 · lost 1
GPT-5won 15 · lost 1
Claude Sonnet 4.5won 15 · lost 1
Grok 3 Miniwon 14 · lost 1
Grok 4won 14 · lost 1
GPT-4.1won 14 · lost 1
GPT-5.6 Lunawon 21 · lost 2
DeepSeek V4 Flashwon 20 · lost 2
GPT-5.2won 17 · lost 2
O3won 14 · lost 2
Claude Opus 4won 13 · lost 2
Qwen3.8 27Bwon 19 · lost 3
Claude Sonnet 5won 19 · lost 3
Grok 4.5won 19 · lost 3
GPT-5.4won 18 · lost 3
GLM-5.2won 18 · lost 3
Qwen3.7 Maxwon 17 · lost 3
Claude Opus 4.5won 16 · lost 3
Mimo V2.6 Flashwon 14 · lost 3
GPT-5.6 Terrawon 18 · lost 4
Grok 4.7won 17 · lost 4
DeepSeek V4 Prowon 17 · lost 4
Qwen3.8 Maxwon 17 · lost 4
Gemini 3 Prowon 12 · lost 3
Claude Sonnet 4.6won 16 · lost 4
Gemini 3.6 Flashwon 18 · lost 5
GLM-5.3 Flashwon 17 · lost 5
Gemini 3.1 Prowon 17 · lost 5
GLM-5.3won 16 · lost 5
Claude Opus 4.7won 16 · lost 5
Muse Spark 1.1won 15 · lost 5
GPT-6 Solwon 17 · lost 6
Claude Opus 4.6won 14 · lost 5
Grok 4.6won 16 · lost 6
DeepSeek V4.1 Flashwon 17 · lost 7
Mimo V2.6 Prowon 11 · lost 5
Gemini 3.5 Flashwon 14 · lost 7
Claude Opus 4.8won 14 · lost 7
Gemini 3.7 Flashwon 15 · lost 8
Muse Spark 1.2won 14 · lost 8
Gemini 3.8 Flashwon 14 · lost 9
GPT-5.5won 12 · lost 10
Kimi K3won 12 · lost 10
GPT-6.1 Solwon 11 · lost 11
GPT-6 Astrawon 12 · lost 12
Claude Opus 5won 10 · lost 13
GPT-5.6 Solwon 9 · lost 15
Muse Spark 1.3won 6 · lost 13
Gemini 4 Argonwon 5 · lost 12
Claude Fable 5won 5 · lost 16
Claude Opus 5.5won 5 · lost 19
Claude Fable 5.1won 4 · lost 19

§ 3 · Sources

Where the numbers come from

8 publications, 30 figures. Every one links to the page it was read from.

LMArena Text ↗1 measureread 2026-10-04
LMArena Text 1471
Artificial Analysis ↗1 measureread 2026-10-04
Artificial Analysis Intelligence Index 56
GDPval-AA ↗1 measureread 2026-10-04
GDPval-AA 1840
AA-Briefcase ↗1 measureread 2026-10-04
AA-Briefcase 1824
LiveBench ↗8 measuresread 2026-10-04
LiveBench 77.8LiveBench · Reasoning 86.8LiveBench · Coding 88.9LiveBench · Agentic Coding 39.3LiveBench · Mathematics 96.7LiveBench · Data Analysis 78.6LiveBench · Language 83.4LiveBench · Instruction Following 70.5
SimpleBench ↗1 measureread 2026-10-04
SimpleBench 75.9%
Vals.ai ↗11 measuresread 2026-10-04
Vals · Legal Research Bench 48.08%Vals · Harvey Legal Agent Benchmark 2.92%Vals · Finance Agent 58.1%Vals · MedCode 52.92%Vals · MedScribe 91.1%Vals · BioMysteryBench 81.11%Vals · CyberBench 59.58%Vals · SRE Bench 30.15%Vals · Vibe Code Bench 92.39% (OpenHands)Vals · Code Migration 69.83%Vals · ProofBench 100%
ParseBench ↗6 measuresread 2026-10-04
ParseBench · Overall 70.18%ParseBench · Tables 91.04%ParseBench · Charts 37.69%ParseBench · Content faithfulness 91.64%ParseBench · Semantic formatting 71.63%ParseBench · Visual grounding 58.92%

Badge

PublicAI Index badge for Claude Sonnet 5.5[![PublicAI Index](https://publicai.io/model-index/badge?model=claude-sonnet-5-5)](https://publicai.io/model-index/m/claude-sonnet-5-5)