‹ PublicAI Index
The LLM benchmark aggregator.
Llama 3.1 Instruct Turbo (70B)
Meta · 70B
Strongest in Safe-prompt compliance (#54 of 82), weakest in Safety (#279 of 337). Above par in 5 of 10 scopes. Among the models it meets almost everywhere, it finishes behind GPT-5.1 and GPT-5 and ahead of Grok 3 and Grok 4.1 Fast.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Professional42.5−19.1#145/1681/1
Medical50.7−16.1#74/1401/1
Finance33.3−29.6#147/1511/1
Safety45.7−15.8#279/3372/3
Safe-prompt compliance50.8−9.6#54/821/1
Jailbreak resistance55.4−9.5#112/2721/1
Fairness50.2−19.8#113/3002/2
Toxicity avoidance52.7−6.3#161/2721/1
Harm refusal43.3−18.3#268/3002/2
Secure code26−40.5#270/2741/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Llama 3.1 Instruct Turbo (70B), left for the other.
§ 3 · Sources
Where the numbers come from
3 publications, 14 figures. Every one links to the page it was read from.
Vals · CorpFin 38.85%Vals · TaxEval 56.17%Vals · MedQA 84.78%
HELM Safety · HarmBench 46.9%HELM Safety · SimpleSafetyTests 92.5%HELM Safety · Anthropic Red Team 93.2%HELM Safety · BBQ 95.4%HELM Safety · XSTest 94.5%
Enkrypt · Jailbreak risk 8.2%Enkrypt · Harmful content risk 37.2%Enkrypt · CBRN risk 13.5%Enkrypt · Toxicity risk 4.4%Enkrypt · Bias risk 86.3%Enkrypt · Insecure code risk 92.9%
Badge
[](https://publicai.io/model-index/m/llama-3-1-instruct-turbo-70b)