‹ PublicAI Index
The LLM benchmark aggregator.
DeepSeek LLM 67B Chat
DeepSeek · 67B
Strongest in Safe-prompt compliance (#72 of 82), weakest in Human preference (#282 of 342). Above par in 1 of 9 scopes. Among the models it meets almost everywhere, it finishes behind GPT-5.1 and GPT-5 and ahead of Command R Plus and GPT-3.5 Turbo.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety46.7−14.8#262/3372/3
Safe-prompt compliance38.4−22#72/821/1
Secure code55.8−10.7#113/2741/1
Jailbreak resistance48−16.9#192/2721/1
Toxicity avoidance49.1−9.9#208/2721/1
Harm refusal46.9−14.7#210/3002/2
Fairness43.7−26.3#232/3002/2
Human preference36.3−31.2#282/3421/1
Human preference36.3−31.2#282/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for DeepSeek LLM 67B Chat, left for the other.
§ 3 · Sources
Where the numbers come from
3 publications, 12 figures. Every one links to the page it was read from.
LMArena Text 1185
HELM Safety · HarmBench 64.9%HELM Safety · SimpleSafetyTests 96.8%HELM Safety · Anthropic Red Team 99.4%HELM Safety · BBQ 86.2%HELM Safety · XSTest 88.9%
Enkrypt · Jailbreak risk 14.5%Enkrypt · Harmful content risk 61.7%Enkrypt · CBRN risk 16.5%Enkrypt · Toxicity risk 6.9%Enkrypt · Bias risk 90.4%Enkrypt · Insecure code risk 21.8%
Badge
[](https://publicai.io/model-index/m/deepseek-llm-67b-chat)