‹ PublicAI Index
The LLM benchmark aggregator.
Claude 3.5 Haiku
Anthropic
Strongest in Jailbreak resistance (#5 of 272), weakest in Secure code (#217 of 274). Above par in 5 of 20 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Gemini 1.5 Pro and Grok Build 0.1.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety58.8−2.7#13/3371/3
Jailbreak resistance63.8−1.1#5/2721/1
Harm refusal60.1−1.5#8/3001/2
Fairness70leads#10/3001/2
Toxicity avoidance58.1−0.9#27/2721/1
Secure code40.7−25.8#217/2741/1
Coding46.5−22.6#105/1651/5
Agentic coding45.7−23#102/1571/4
Knowledge32.2−31.3#133/1381/2
Academic knowledge28.6−35.3#117/1231/1
Professional43.3−18.3#140/1681/1
Legal44−23.1#110/1511/1
Finance39.6−23.3#131/1511/1
Reasoning33.8−33.1#175/1781/4
Science26.1−37.4#116/1221/1
Mathematics33.3−32.4#135/1401/2
Human preference49.8−17.7#194/3421/1
Human preference49.8−17.7#194/3421/1
Agents44.1−24#199/2681/5
Knowledge work41.1−32.5#127/1781/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Claude 3.5 Haiku, left for the other.
§ 3 · Sources
Where the numbers come from
5 publications, 15 figures. Every one links to the page it was read from.
LMArena Text 1325
GDPval-AA 198
Aider polyglot 28%
Vals · LegalBench 70.33%Vals · CorpFin 50.82%Vals · TaxEval 57.36%Vals · GPQA Diamond 37.88%Vals · MMLU Pro 64.12%Vals · AIME 3.33%
Enkrypt · Jailbreak risk 1.1%Enkrypt · Harmful content risk 5.6%Enkrypt · CBRN risk 3%Enkrypt · Toxicity risk 0.6%Enkrypt · Bias risk 43.4%Enkrypt · Insecure code risk 52.4%
Badge
[](https://publicai.io/model-index/m/claude-3-5-haiku)