‹ PublicAI Index
The LLM benchmark aggregator.
Claude 3.7 Sonnet
Anthropic
Strongest in Jailbreak resistance (#3 of 272), weakest in Secure code (#224 of 274). Above par in 15 of 23 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Grok 4.1 Fast and GPT-5.4 Mini.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Coding54.4−14.7#36/1651/5
Agentic coding55.3−13.4#37/1571/4
Professional53.7−7.9#46/1681/1
Finance54.7−8.2#49/1511/1
Legal53.3−13.8#52/1511/1
Medical53.6−13.2#53/1401/1
Safety55.3−6.2#50/3372/3
Jailbreak resistance64.3−0.6#3/2721/1
Harm refusal59−2.6#17/3002/2
Toxicity avoidance58−1#33/2721/1
Safe-prompt compliance55.1−5.3#35/821/1
Fairness49.3−20.7#125/3002/2
Secure code38.1−28.4#224/2741/1
Knowledge51.6−11.9#79/1381/2
Academic knowledge51.9−12#73/1231/1
Reasoning47.1−19.8#109/1782/4
Reasoning48.3−20.8#69/1401/3
Science49.4−14.1#78/1221/1
Mathematics44.2−21.5#99/1401/2
Human preference55.9−11.6#135/3421/1
Human preference55.9−11.6#135/3421/1
Agents46.9−21.2#154/2681/5
Computer use44.5−27.2#24/371/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Claude 3.7 Sonnet, left for the other.
§ 3 · Sources
Where the numbers come from
7 publications, 23 figures. Every one links to the page it was read from.
LMArena Text 1388
Aider polyglot 64.9%
OSWorld 35.8%
SimpleBench 46.4%
Vals · LegalBench 82.52%Vals · CorpFin 60.41%Vals · TaxEval 74.04%Vals · MortgageTax 66.85%Vals · MedQA 90.22%Vals · GPQA Diamond 75.25%Vals · MMLU Pro 82.73%Vals · AIME 44.58%
HELM Safety · HarmBench 84.3%HELM Safety · SimpleSafetyTests 100%HELM Safety · Anthropic Red Team 99.7%HELM Safety · BBQ 92.1%HELM Safety · XSTest 96.4%
Enkrypt · Jailbreak risk 0.7%Enkrypt · Harmful content risk 1.7%Enkrypt · CBRN risk 4.8%Enkrypt · Toxicity risk 0.7%Enkrypt · Bias risk 84%Enkrypt · Insecure code risk 57.8%
Badge
[](https://publicai.io/model-index/m/claude-3-7-sonnet)