‹ PublicAI Index
The LLM benchmark aggregator.
Llama 3.1 Instruct Turbo (405B)
Meta · 405B
Strongest in Safe-prompt compliance (#41 of 82), weakest in Secure code (#239 of 274). Above par in 6 of 10 scopes. Among the models it meets almost everywhere, it finishes behind GPT-5 and Claude Sonnet 4.5 and ahead of Gemini 3.6 Flash and GLM-5.2.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Professional48.5−13.1#96/1681/1
Medical52.5−14.3#65/1401/1
Finance43.6−19.3#120/1511/1
Safety50.1−11.4#180/3372/3
Safe-prompt compliance53.9−6.5#41/821/1
Jailbreak resistance58.9−6#60/2721/1
Fairness49.4−20.6#121/3002/2
Harm refusal50.6−11#145/3002/2
Toxicity avoidance53.3−5.7#148/2721/1
Secure code33.2−33.3#239/2741/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Llama 3.1 Instruct Turbo (405B), left for the other.
§ 3 · Sources
Where the numbers come from
3 publications, 13 figures. Every one links to the page it was read from.
Vals · TaxEval 60.88%Vals · MedQA 88.24%
HELM Safety · HarmBench 62.7%HELM Safety · SimpleSafetyTests 98.8%HELM Safety · Anthropic Red Team 96.5%HELM Safety · BBQ 94.5%HELM Safety · XSTest 95.9%
Enkrypt · Jailbreak risk 5.3%Enkrypt · Harmful content risk 26.7%Enkrypt · CBRN risk 9.8%Enkrypt · Toxicity risk 4%Enkrypt · Bias risk 87.1%Enkrypt · Insecure code risk 67.6%
Badge
[](https://publicai.io/model-index/m/llama-3-1-instruct-turbo-405b)