‹ PublicAI Index
The LLM benchmark aggregator.
Llama 3.1 Tulu 3 8B SFT
Ai2 · 8B
Strongest in Harm refusal (#68 of 299, on 1 of its 2 boards), weakest in Fairness (#262 of 299). Above par in 5 of 6 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 5 and MiniMax M2.7 and ahead of Nova Micro and K2 Chat.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety52.8−8.7#107/3361/3
Harm refusal55.2−6.4#68/2991/2
Jailbreak resistance57.6−7.3#76/2721/1
Toxicity avoidance55.3−3.7#101/2721/1
Secure code53.6−12.9#125/2741/1
Fairness42−28#262/2991/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Llama 3.1 Tulu 3 8B SFT, left for the other.
§ 3 · Sources
Where the numbers come from
1 publication, 6 figures. Every one links to the page it was read from.
Enkrypt · Jailbreak risk 6.4%Enkrypt · Harmful content risk 13.9%Enkrypt · CBRN risk 11.5%Enkrypt · Toxicity risk 2.6%Enkrypt · Bias risk 90.2%Enkrypt · Insecure code risk 26.2%
Badge
[](https://publicai.io/model-index/m/llama-3-1-tulu-3-8b-sft)