‹ PublicAI Index
The LLM benchmark aggregator.
Claude Haiku 4.5
Anthropic
Strongest in Tool use (#6 of 81), weakest in Secure code (#194 of 274). Above par in 11 of 27 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of DeepSeek R1 and Grok 3.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety55.4−6.1#44/3373/3
Jailbreak resistance61.7−3.2#26/2721/1
Harm refusal57.6−4#27/3002/2
Toxicity avoidance57.4−1.6#42/2721/1
Fairness57−13#50/3002/2
Factual grounding50.4−20.2#56/1011/1
Safe-prompt compliance47.9−12.5#61/821/1
Secure code45.1−21.4#194/2741/1
Agents55−13.1#65/2683/5
Tool use72.4−1.6#6/811/1
Knowledge work47.3−26.3#94/1782/2
Reasoning49.1−17.8#94/1782/4
Mathematics54.4−11.3#59/1401/2
Science47.3−16.2#84/1221/1
Reasoning43.6−25.5#89/1401/3
Knowledge47.4−16.1#101/1381/2
Academic knowledge46.9−17#93/1231/1
Human preference58.4−9.1#109/3421/1
Human preference58.4−9.1#109/3421/1
Professional46−15.6#121/1681/1
Medical47.2−19.6#95/1401/1
Finance46.3−16.6#106/1511/1
Legal44.2−22.9#107/1511/1
Core abilities46.3−21.1#135/2041/3
General intelligence44.7−25.4#137/2041/3
Coding41.9−27.2#143/1651/5
Agentic coding40.9−27.8#137/1571/4
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Claude Haiku 4.5, left for the other.
§ 3 · Sources
Where the numbers come from
10 publications, 35 figures. Every one links to the page it was read from.
LMArena Text 1414
Artificial Analysis Intelligence Index 17
GDPval-AA 719
AA-Briefcase 616
ARC-AGI-2 4%
BFCL v4 68.7%
Vals · Legal Research Bench 10.58%Vals · CaseLaw 56.48%Vals · LegalBench 81.24%Vals · Harvey Legal Agent Benchmark 0.83%Vals · Finance Agent 31.01%Vals · CorpFin 60.61%Vals · TaxEval 67.54%Vals · MortgageTax 62.16%Vals · MedQA 79.57%Vals · MedCode 32.68%Vals · MedScribe 85.23%Vals · SWE-bench Verified 66.6% (Mini-SWE-agent)Vals · Vibe Code Bench 11.39% (OpenHands)Vals · Code Migration 7.55%Vals · GPQA Diamond 72.22%Vals · MMLU Pro 78.72%Vals · AIME 82.71%
HELM Safety · HarmBench 95.9%HELM Safety · SimpleSafetyTests 98.8%HELM Safety · Anthropic Red Team 96.9%HELM Safety · BBQ 92.8%HELM Safety · XSTest 93.2%
Enkrypt · Jailbreak risk 2.9%Enkrypt · Harmful content risk 1.1%Enkrypt · CBRN risk 10.5%Enkrypt · Toxicity risk 1.1%Enkrypt · Bias risk 64.9%Enkrypt · Insecure code risk 43.6%
Vectara · Factual consistency 90.2%
Badge
[](https://publicai.io/model-index/m/claude-haiku-4-5)