‹ PublicAI Index
The LLM benchmark aggregator.
Llama 3 8B Instruct
Meta · 8B
Strongest in Safe-prompt compliance (#45 of 82), weakest in Human preference (#267 of 342). Above par in 6 of 9 scopes. Among the models it meets almost everywhere, it finishes behind GPT-5.1 and GPT-5 and ahead of DeepSeek V3 and Command R Plus.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety51.2−10.3#154/3372/3
Safe-prompt compliance53.3−7.1#45/821/1
Harm refusal53.7−7.9#94/3002/2
Secure code53.4−13.1#126/2741/1
Jailbreak resistance51.6−13.3#164/2721/1
Toxicity avoidance51.6−7.4#182/2721/1
Fairness42.4−27.6#256/3002/2
Human preference40.1−27.4#267/3421/1
Human preference40.1−27.4#267/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Llama 3 8B Instruct, left for the other.
§ 3 · Sources
Where the numbers come from
3 publications, 12 figures. Every one links to the page it was read from.
LMArena Text 1224
HELM Safety · HarmBench 72.7%HELM Safety · SimpleSafetyTests 99.3%HELM Safety · Anthropic Red Team 98.8%HELM Safety · BBQ 76.5%HELM Safety · XSTest 95.6%
Enkrypt · Jailbreak risk 11.4%Enkrypt · Harmful content risk 12.8%Enkrypt · CBRN risk 14.7%Enkrypt · Toxicity risk 5.2%Enkrypt · Bias risk 80.4%Enkrypt · Insecure code risk 26.7%
Badge
[](https://publicai.io/model-index/m/llama-3-8b-instruct)