‹ PublicAI Index
The LLM benchmark aggregator.
DeepSeek R1
DeepSeek · 685B · open weights
Strongest in Safe-prompt compliance (#2 of 82), weakest in Safety (#285 of 337). Above par in 14 of 26 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of GPT OSS 120B and Grok 4.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Coding57.7−11.4#18/1652/5
Code generation59−11#15/771/2
Agentic coding57−11.7#29/1571/4
Core abilities55.7−11.7#39/2041/3
General intelligence58.2−11.9#39/2041/3
Knowledge52−11.5#76/1381/2
Academic knowledge52.4−11.5#70/1231/1
Professional48.3−13.3#98/1681/1
Medical53.9−12.9#51/1401/1
Finance49.1−13.8#92/1511/1
Legal41.7−25.4#126/1511/1
Human preference59.1−8.4#102/3421/1
Human preference59.1−8.4#102/3421/1
Reasoning47.1−19.8#107/1783/4
Mathematics52−13.7#71/1401/2
Reasoning43−26.1#95/1402/3
Agents39.8−28.3#240/2682/5
Knowledge work36.8−36.8#150/1782/2
Safety45.3−16.2#285/3373/3
Safe-prompt compliance60.4leads#2/821/1
Fairness55.1−14.9#67/3002/2
Factual grounding46.7−23.9#71/1011/1
Toxicity avoidance56−3#77/2721/1
Jailbreak resistance32.9−32#247/2721/1
Secure code30.2−36.3#251/2741/1
Harm refusal41.4−20.2#281/3002/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for DeepSeek R1, left for the other.
§ 3 · Sources
Where the numbers come from
12 publications, 26 figures. Every one links to the page it was read from.
LMArena Text 1422
GDPval-AA 282
AA-Briefcase 130
ARC-AGI-2 1.1%
LiveCodeBench 73.1%
Aider polyglot 71.4%
Kagi LLM Benchmark 69.4%
SimpleBench 40.8%
Vals · LegalBench 67.32%Vals · CorpFin 54.12%Vals · TaxEval 72.28%Vals · MedQA 90.8%Vals · MMLU Pro 83.18%Vals · AIME 73.96%
HELM Safety · HarmBench 54.6%HELM Safety · SimpleSafetyTests 98.3%HELM Safety · Anthropic Red Team 99.1%HELM Safety · BBQ 96.5%HELM Safety · XSTest 98.8%
Enkrypt · Jailbreak risk 27.2%Enkrypt · Harmful content risk 57.8%Enkrypt · CBRN risk 53.5%Enkrypt · Toxicity risk 2.1%Enkrypt · Bias risk 74.9%Enkrypt · Insecure code risk 73.8%
Vectara · Factual consistency 88.7%
Badge
[](https://publicai.io/model-index/m/deepseek-r1)