‹ PublicAI Index
The LLM benchmark aggregator.
Llama 3.2 3B Instruct
Meta · 3.2B · open weights
Strongest in Tool use (#71 of 81), weakest in Human preference (#296 of 342). Above par in 3 of 10 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 5 and Claude Sonnet 4.5 and ahead of DeepSeek R1 and DeepSeek V3.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents43.6−24.5#205/2681/5
Tool use38.5−35.5#71/811/1
Safety49−12.5#209/3371/3
Harm refusal53.9−7.7#90/3001/2
Jailbreak resistance55.7−9.2#106/2721/1
Toxicity avoidance52.6−6.4#164/2721/1
Secure code32.6−33.9#241/2741/1
Fairness42.9−27.1#247/3001/2
Human preference34.6−32.9#296/3421/1
Human preference34.6−32.9#296/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Llama 3.2 3B Instruct, left for the other.
§ 3 · Sources
Where the numbers come from
3 publications, 8 figures. Every one links to the page it was read from.
LMArena Text 1167
BFCL v4 21.95%
Enkrypt · Jailbreak risk 8%Enkrypt · Harmful content risk 15%Enkrypt · CBRN risk 14.3%Enkrypt · Toxicity risk 4.5%Enkrypt · Bias risk 88.9%Enkrypt · Insecure code risk 68.9%
Badge
[](https://publicai.io/model-index/m/llama-3-2-3b-instruct)